
Have you ever encountered the following issues while working on text annotation or data labeling tasks?
• Annotators’ judgments differ slightly
• There is no consistency or stability in how labels are assigned
• Model accuracy does not improve as expected
The cause of these issues often lies not only in the individual skills of annotators but also in the "guideline design." The quality of the guidelines greatly affects the quality of the training data, and consequently, the accuracy of NLP (Natural Language Processing) and LLM (Large Language Model) models.
This article focuses on the design of guidelines for text annotation and organizes and explains the concepts necessary to achieve highly reproducible training data creation and data labeling.
• Why guidelines are important
• What content should be defined
• Common practical mistakes and countermeasures
- Table of Contents
-
- 1. What Are Guidelines in Annotation?
- 2. Why Is Guideline Design Important in Text Annotation?
- 3. Main Items to Define in Guidelines
- 3-1. Label Definitions
- 3-2. Criteria and Prioritization
- 3-3. Presentation of Concrete Examples and NG Cases
- 4. How to Proceed with Guideline Design (Practical Workflow)
- 4-1. Common Failures in Guideline Design
- 4-2. Conditions for Guidelines Leading to High-Quality Annotation
- 5. Summary: The Quality of Text Annotation Is Determined by Guideline Design
- 6. Human Science Support for Text Annotation
1. What Are Guidelines in Annotation?

The guidelines used in annotation are documents that verbalize and compile rules on "by what criteria and how to interpret certain data to assign labels."
It is often confused with a manual that summarizes work procedures, but essentially it is different. What the guidelines define are the very criteria that the model should learn to make judgments.
• What is considered to have the same meaning
• What to focus on when making a judgment
• Which to consider correct when in doubt
Especially in NLP and LLM development, how concretely and explicitly the implicit judgments made by humans can be formalized into rules greatly affects data quality.
2. Why Is Guideline Design Important in Text Annotation?

Variation in judgment = Variation in training data = Variation in model quality
Text data has the following characteristics.
• Context-dependent
• Many polysemous words
• Large variation in expressions
Example of context dependence:
“This is nice.”
→ Without knowing the subject just mentioned (product, document, proposal, etc.), the meaning cannot be determined.
Example of many polysemous words:
“Hashi” → bridge (structure), chopsticks (tableware), edge (end), etc.
Example of large variation in expressions:
Words with multiple expressions for the same meaning (such as “cancellation,” “cancel,” “call off”) and distinctions in kanji, katakana, and hiragana like “dog,” “Inu,” and “inu”
If annotation proceeds with ambiguous guidelines, differences in interpretation among annotators are likely to occur. These interpretation differences directly become noise in the labels that the model learns from, leading to decreased output accuracy and stability.
Places where humans hesitate are also prone to model mislearning.
Sections where annotators "struggle to decide" are also parts where the model finds it difficult to learn the correct answer. That is precisely why anticipating cases where people might hesitate and proactively establishing rules is the most important point in guideline design.
Related blog: 7 Tips for Successful Annotation
3. Main Items to Define in Guidelines

From here, we will explain the main items that should be defined in the guidelines.
3-1. Label Definitions
The first important point is to clearly state the definition of each label in writing.
• Meaning of the label
• Scope of application
• Differences from similar labels
Additionally, by clearly specifying cases where labels should not be assigned and examples where boundaries tend to be ambiguous, it is possible to reduce variations in interpretation.
3-2. Criteria and Prioritization
In text annotation, cases allowing multiple interpretations frequently occur.
Therefore, it is important to establish "guidelines for when in doubt."
• Handling cases where context is insufficient
• Priority order when there are multiple label candidates within a single sentence
• Decision rules when the evidence is weak
3-3. Presentation of Concrete Examples and NG Cases
It is difficult to fully share judgment criteria with definitions alone.
What is effective in such cases is the presentation of concrete examples and bad examples.
• Correct Examples (Good Examples)
• Common Mistake Examples (Bad Examples)
• Edge Cases
In practice, the more examples there are, the higher the resolution of the guidelines tends to be.
4. How to Proceed with Guideline Design (Practical Workflow)

Guideline design generally proceeds as follows.
• Clarify the purpose of annotation
• Design the label system
• Create the initial guideline
• Trial annotation
• Reflect feedback and revise
• Start full-scale operation
The important thing is not to consider it complete once it is created.
It is necessary to collect and review questions, comments, and cases where judgment was difficult that arise during the trial phase and actual work, and reflect them in the guidelines to develop them into more practical and comprehensive ones.
4-1. Common Failures in Guideline Design
Common failures include the following.
• Definitions are too abstract, allowing for a wide range of interpretations
• Few concrete examples, making the judgment criteria unclear
• Rules change midway through the project
• Guidelines are not updated
• The latest version of the guidelines is not shared
The latter three points in particular are easy to overlook during actual work.
Not only rule changes due to client requests but also standard changes from reinterpretations of guidelines and edge cases tend to be shared only verbally, by email, or chat while being pressed by daily annotation tasks.
As a result, updates to the guidelines are postponed, leaving outdated information in documents that should be current, and the latest rules are not sufficiently shared on-site.
These issues not only lead to a decline in quality but also cause rework. Consequently, they result in increased costs and schedule delays.
4-2. Conditions for Guidelines Leading to High-Quality Annotation
High-quality guidelines share the following common characteristics.
• Anyone reading it can make the same judgment
• Exceptions and easily confusing cases are organized
• Update history is managed
• Linked with the quality control process
It is important to view guidelines not as "work manuals" but as blueprints that support quality.
Additionally, providing opportunities for annotators to discuss the guidelines among themselves is indispensable for good guideline design. Having such forums prevents delays in sharing changes to rules and standards, making it easier to align understanding among annotators. Furthermore, by sharing and resolving edge cases and questions that arise during actual work, the guidelines can be enhanced, leading to higher-quality annotations.
5. Summary: The Quality of Text Annotation Is Determined by Guideline Design
The quality of text annotation is greatly influenced by guideline design. That is why it is important to properly design the guidelines and continuously improve them while operating. To succeed in creating training data and data labeling, it is crucial to clarify the decision criteria before starting annotation and to operate while maintaining them, as this greatly affects subsequent quality and efficiency.
6. Human Science Support for Text Annotation
Over 48 million pieces of training data created
At Human Science, we participate in AI model development projects across a wide range of industries, starting with natural language processing and extending to medical support, automotive, IT, manufacturing, and construction. We have a proven track record of direct business dealings with many companies, including GAFAM. Additionally, we have provided over 48 million pieces of high-quality training data for AI development projects. From small-scale projects to large long-term projects with a team of 150 annotators, we handle various types of training data creation, data labeling, and data structuring regardless of the industry.
Resource management without crowdsourcing
At Human Science, we do not use crowdsourcing. Instead, projects are handled by personnel who are contracted with us directly. Based on a solid understanding of each member's practical experience and their evaluations from previous projects, we form teams that can deliver maximum performance.
Supports not only curation and annotation but also the creation and structuring of generative AI LLM datasets
In addition to labeling for data organization and annotation for identification-based AI systems, we also support the structuring of document data for generative AI and LLM RAG construction. Since our founding, we have been engaged in manual production as a main business and service, and we provide optimal solutions leveraging our unique know-how and deep understanding of various document structures.
Secure room available on-site
Within our Shinjuku office at Human Science, we have secure rooms that meet ISMS standards. Therefore, we can guarantee security, even for projects that include highly confidential data. We consider the preservation of confidentiality to be extremely important for all projects. When working remotely as well, our information security management system has received high praise from clients, because not only do we implement hardware measures, we continuously provide security training to our personnel.
Reference link: Text Annotation Service

Text Annotation
Audio Annotation
Image & Video Annotation
Generative AI, LLM, RAG Data Structuring
AI Model Development
In-House Support
For the medical industry
For the automotive industry
For the IT industry
For the manufacturing industry































































































