
- Table of Contents
-
- 1. Introduction
- 2. In Conventional AI Development, the Quality of "Training Data" Was Central
- 3. In the Era of Generative AI, the Quality of "Reference Data" Becomes Important
- 4. The Output of Generative AI is Verified for Quality Using "Evaluation Data"
- 5. In Agent AI, Organizing Data and Business Rules Becomes Crucial
- 6. What Constitutes "Good Data" Going Forward
- 7. Human Science Teacher Data Creation, LLM RAG Data Structuring Outsourcing Service
1. Introduction

In recent years, with the spread of generative AI such as ChatGPT, AI has been rapidly permeating our work and daily lives. Furthermore, in recent years, AI agents that autonomously execute multiple tasks have also attracted significant attention.
Our company is also working on business improvement and verification using AI agents, and we have increasingly experienced their effectiveness in supporting routine tasks that occur regularly, such as information gathering and invoice processing. On the other hand, even if the performance of AI models improves, it does not necessarily lead to better business outcomes by itself. When using AI in business, it is necessary to verify which information the AI refers to, what kind of output is considered usable, and how to distinguish incorrect information or unnatural responses.
What becomes important at this time is "data quality." However, in the era of generative AI and agent AI, data quality is not something that can be ensured by looking only at traditional training data. It is necessary to separately consider the data used to train the AI, the data referenced by the AI, and the data used to evaluate the AI's output.
This article reflects on the quality of training data emphasized in conventional AI development and considers what constitutes "good data" in the era of generative AI and agent AI.
[Reference Blogs]
>Why Domain-Specific LLMs Don’t Get "Smarter" — Three Points That Hinder Quality Improvement
>What Is the Evaluation of Generative AI? Explaining How to Measure the Quality of LLMs
2. In Conventional AI Development, the Quality of "Training Data" Was Central

In conventional AI development, AI systems with relatively clear tasks, such as image recognition and text classification, have been widely utilized. For image recognition AI, labels like "dog," "cat," and "car" are attached to images; in object detection, the position of the target object is enclosed in a box; and in segmentation, the regions are specified in detail.
In such annotations, it is important to first decide "what is considered correct." When the object is only partially visible or overlaps with the background making it hard to see, if the criteria remain ambiguous, the way labels are assigned can vary from one annotator to another.
Therefore, in annotation work, it is important to establish work rules, reduce variability among workers, and perform checks and corrections as needed. The "good data" of this era can be said to be training data that is not only correctly labeled but also created with stable quality according to certain rules.
【Reference Blogs】
>Where Do Failures in AI Development for Manufacturing Originate? — Tips for Success from the Perspective of Data Quality
>【Spin-off】Management of Annotation Work in Human Science — Haste Makes Waste for Quality Assurance. Taking the Long Way Around Becomes the Shortcut
>【Spin-off】Does Repeated Checking Improve Annotation Quality? — Quality Control at Our Annotation Site
>How to Ensure and Improve the Quality of Training Data? Explanation of Practical Methods!
3. In the Era of Generative AI, the Quality of "Reference Data" Becomes Important

In the era of generative AI, the concept of data quality can no longer be confined to just training data. Generative AI creates text or answers questions based on a vast amount of information, but if the source information is outdated, incorrect, or lacks context, the reliability of the output also decreases.
Additionally, with the spread of generative AI, the amount of AI-generated text and images is also increasing. As low-quality AI-generated content grows, useful information tends to get buried, making it necessary to discern which information can be trusted.
When companies utilize generative AI in their operations, relying solely on general information available on the internet is often insufficient. There are situations where data not publicly disclosed, such as product specifications, operation manuals, internal rules, and past inquiry records, become crucial. Even when using internal generative AI or RAG, if the source documents are not well organized, the expected responses cannot be obtained.
In other words, in the era of generative AI, "good data" includes not only data for AI training but also highly reliable business data for AI reference. Organizing which materials should be considered as correct evidence and which information AI should refer to becomes a prerequisite for AI utilization.
[Reference Blog]
>Essential for LLM Development! Key Points for Translation Data Evaluation and High-Quality Data Preparation
4. The Output of Generative AI is Verified for Quality Using "Evaluation Data"

Even if the training data and reference data are well prepared, it does not necessarily mean that AI output can be used directly in business operations. Many answers generated by AI do not have a single correct solution. When multiple answers are possible for the same question, determining which is most appropriate cannot be judged by simple right or wrong.
Evaluating the output of generative AI requires multiple perspectives, such as whether the content is based on facts, aligns with the intent of the question, is easy for the reader to understand, and poses no issues from the standpoint of safety and compliance. Even if the answer appears natural, hallucinations with incorrect content have not yet been completely eliminated.
If the effort required for fact-checking and regeneration becomes too great, the AI introduced to improve efficiency may instead increase the operational burden. That is why it is important to evaluate the output of generative AI and establish criteria to determine which responses can be used in business operations.
【Reference Blog】
>What Is Evaluation of Generative AI? Explaining How to Measure the Quality of LLMs
>The Role of RLHF in Domestic LLMs — Where Does "Human Judgment" That Determines the Quality of Japanese LLMs Come Into Play?
5. In Agent AI, Organizing Data and Business Rules Becomes Crucial

An AI agent is an AI that autonomously carries out multiple tasks according to user instructions. It is expected to be used in various ways, such as information retrieval, comparison of multiple information sources, organizing content that meets certain conditions, and supporting regular administrative tasks.
With such AI agents, it may not be sufficient to evaluate them solely based on the final output. Even if the investigation results are organized, if the referenced information is outdated or the basis is unclear, it cannot be used directly in business operations.
Even when AI supports routine tasks such as processing invoices, it is not enough to simply verify whether the amounts and client names have been correctly read. It is also necessary to check for inconsistencies in the invoice date or payment deadline, significant discrepancies compared to past invoices, whether the amount requires approval, and whether the document can be forwarded to the next step in the workflow as is.
To utilize agent AI in business operations, it is necessary to organize not only the quality of reference data but also the business rules such as under what conditions processing can proceed and from where human verification is required. Here, internal rules and judgment criteria themselves must also be treated as part of the data supporting AI utilization.
[Reference Blog]
>What AI Agents Can Do: "AI That Thinks and Acts on Its Own" Will Change How We Work
6. What Constitutes "Good Data" Going Forward

As we have seen so far, there is not just one type of "good data" in the era of generative AI and agent AI. Traditional AI development mainly focused on training data created with correct labels and consistent annotations.
On the other hand, when utilizing generative AI in business operations, the quality of internal documents and business knowledge referenced by the AI becomes crucial. Furthermore, standards for evaluating the AI’s output responses and decisions are indispensable. In agent AI, it is also necessary to establish business rules regarding which information to use, the procedures for processing, and where human verification is required.
In other words, the "good data" of the future can be said to involve separating training data, reference data, and evaluation data, and organizing each according to the AI's intended use. Furthermore, even as AI evolves, it will be up to us humans to decide what to teach, what to reference, and which outputs to judge as good.
[Reference Blog]
>Annotation in the Era of Generative AI: Areas That Can and Cannot Be Automated
7. Human Science Teacher Data Creation, LLM RAG Data Structuring Outsourcing Service
Over 48 million pieces of training data created
At Human Science, we are involved in AI model development projects across various industries, starting with natural language processing, including medical support, automotive, IT, manufacturing, and construction. Through direct transactions with many companies, including GAFAM, we have provided over 48 million high-quality training data. We handle a wide range of training data creation, data labeling, and data structuring, from small-scale projects to long-term large projects with a team of 150 annotators, regardless of the industry.
Resource management without crowdsourcing
At Human Science, we do not use crowdsourcing. Instead, projects are handled by personnel who are contracted with us directly. Based on a solid understanding of each member's practical experience and their evaluations from previous projects, we form teams that can deliver maximum performance.
Generative AI LLM Dataset Creation and Structuring, Also Supporting "Manual Creation and Maintenance Optimized for AI"
We support not only labeling for data organization and training data creation for identification-based AI, but also the structuring of document data for generative AI and LLM RAG construction. Since our founding, manual production has been our main business and service, and we now also provide support for "organizing business knowledge and manualization toward future generative AI and RAG introduction and utilization." We offer optimal solutions leveraging our unique expertise deeply familiar with the structure of various documents.
Secure room available on-site
Within our Shinjuku office at Human Science, we have secure rooms that meet ISMS standards. Therefore, we can guarantee security, even for projects that include highly confidential data. We consider the preservation of confidentiality to be extremely important for all projects. When working remotely as well, our information security management system has received high praise from clients, because not only do we implement hardware measures, we continuously provide security training to our personnel.
In-house Support
We provide staffing services for annotation-experienced personnel and project managers tailored to your tasks and situation. It is also possible to organize a team stationed at your site. Additionally, we support the training of your operators and project managers, assist in selecting tools suited to your circumstances, and help build optimal processes such as automation and work methods to improve quality and productivity. We are here to support your challenges related to annotation and data labeling.

Text Annotation
Audio Annotation
Image & Video Annotation
Generative AI, LLM, RAG Data Structuring
AI Model Development
In-House Support
For the medical industry
For the automotive industry
For the IT industry
For the manufacturing industry































































































