SemEval Shared Task Annotation Designs
Recreations of the annotation designs behind 100 SemEval shared tasks (2013-2026) as runnable Potato configurations. Each folder reproduces the human annotation protocol that produced that task’s dataset. See the documentation to run them.
2013 (1 task)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 9 | Drug-Drug Interaction Extraction from Biomedical Texts | span, radio | link |
2014 (1 task)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 4 | Aspect-Based Sentiment Analysis (Original ABSA) | span, radio | link |
2015 (2 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 4 | TimeLine: Cross-Document Event Ordering | span, radio | link |
| Task 14 | Analysis of Clinical Text: Disorder Identification and Normalization | span, radio | link |
2016 (7 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 2 | Interpretable Semantic Textual Similarity | span, likert | link |
| Task 5 | Aspect-Based Sentiment Analysis | span, radio | link |
| Task 6 | Detecting Stance in Tweets | radio | link |
| Task 7 | Determining Sentiment Intensity of English and Arabic Phrases | radio | link |
| Task 10 | Detecting Minimal Semantic Units and Their Meanings (DiMSUM) | span | link |
| Task 11 | Complex Word Identification | radio | link |
| Task 12 | Clinical TempEval - Temporal Information Extraction from Clinical Notes | span, radio | link |
2017 (5 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 2 | Multilingual Semantic Word Similarity | likert, slider | link |
| Task 3 | Community Question Answering | radio | link |
| Task 5 | Fine-Grained Sentiment Analysis on Financial Microblogs and News | likert, radio | link |
| Task 6 | #HashtagWars - Learning a Sense of Humor | radio | link |
| Task 7 | Detection and Interpretation of English Puns | radio, span | link |
2018 (10 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 2 | Multilingual Emoji Prediction | radio | link |
| Task 3 | Irony Detection in English Tweets | radio | link |
| Task 4 | Character Identification on Multiparty Dialogues | span, radio | link |
| Task 5 | Counting Events and Participants in News | number, text | link |
| Task 6 | Parsing Time Normalizations | span, text | link |
| Task 7 | Semantic Relation Extraction and Classification in Scientific Papers | span, radio | link |
| Task 8 | SecureNLP - Malware Report Semantic Extraction | span, radio, multiselect | link |
| Task 9 | Hypernym Discovery | text, radio | link |
| Task 10 | Capturing Discriminative Attributes | radio | link |
| Task 11 | Machine Comprehension Using Commonsense Knowledge | radio, text | link |
2019 (7 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 3 | EmoContext - Contextual Emotion Detection in Text | radio | link |
| Task 4 | Hyperpartisan News Detection | radio | link |
| Task 5 | HatEval - Multilingual Detection of Hate Speech Against Immigrants and Women | radio, multiselect | link |
| Task 8 | Fact-Checking in Community Question Answering Forums | radio | link |
| Task 9 | Suggestion Mining from Online Reviews and Forums | radio | link |
| Task 10 | Math Question Answering and Category Classification | text, radio | link |
| Task 12 | SemEval-2019 Task 12: Toponym Resolution in Scientific Papers | span, text | link |
2020 (9 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 2 | Predicting Multilingual and Cross-Lingual Lexical Entailment | radio | link |
| Task 3 | Graded Word Similarity in Context | likert, slider | link |
| Task 4 | Commonsense Validation and Explanation | radio, text | link |
| Task 5 | Counterfactual Detection and Reasoning | radio, span | link |
| Task 6 | DeftEval - Extracting Definitions from Free Text | span, radio | link |
| Task 7 | Assessing Humor in Edited News Headlines | likert, radio | link |
| Task 8 | Memotion Analysis - Sentiment and Type Classification of Memes | radio, multiselect | link |
| Task 9 | SemEval-2020 Task 9: Code-Mixed Sentiment (SentiMix) | radio | link |
| Task 10 | Emphasis Selection for Written Text | span | link |
2021 (9 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 1 | Lexical Complexity Prediction | likert, slider | link |
| Task 2 | Multilingual Word-in-Context | radio | link |
| Task 3 | Multilingual Semantic Role Labeling | span, radio | link |
| Task 4 | ReCAM - Reading Comprehension of Abstract Meaning | radio | link |
| Task 6 | SemEval-2021 Task 6: Persuasion Techniques in Memes | multiselect, span | link |
| Task 7 | SemEval-2021 Task 7: HaHackathon Humor Detection | radio, likert | link |
| Task 8 | MeasEval - Counts and Measurements | span, radio, text | link |
| Task 9 | Statement Verification and Evidence Finding with Tables | radio, span | link |
| Task 11 | NLPContributionGraph - Structured Extraction of NLP Contributions | span, text | link |
2022 (10 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 1 | CODWOE - Comparing Dictionaries and Word Embeddings | text, radio | link |
| Task 2 | Idiomaticity Detection | radio | link |
| Task 3 | PreTENS - Presuppositional Acceptability | radio, likert | link |
| Task 4 | Patronizing and Condescending Language Detection | radio, span | link |
| Task 5 | SemEval-2022 Task 5: Multimedia Misogyny (MAMI) | radio, multiselect | link |
| Task 6 | iSarcasmEval: Intended Sarcasm Detection | radio, multiselect, text | link |
| Task 7 | Plausible Clarifications of Implicit and Underspecified Instructions | radio | link |
| Task 8 | Multilingual News Article Similarity | likert | link |
| Task 9 | R2VQ - Recipe Question Answering | text, radio | link |
| Task 12 | Symlink: Linking Mathematical Symbols to their Descriptions | span, radio, text | link |
2023 (10 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 1 | V-WSD - Visual Word Sense Disambiguation | radio | link |
| Task 3 | Detecting Persuasion Techniques in News | multiselect, span | link |
| Task 4 | ValueEval - Human Values behind Arguments | multiselect, radio | link |
| Task 5 | Clickbait Spoiling | text, radio, span | link |
| Task 6 | LegalEval - Legal Document Analysis | radio, span | link |
| Task 7 | Clinical Trial NLI | radio, text | link |
| Task 8 | Causal Medical Claim Identification and PIO Frame Extraction | span | link |
| Task 9 | Tweet Intimacy Analysis | likert | link |
| Task 10 | Explainable Online Sexism Detection | radio, span | link |
| Task 12 | AfriSenti - African Language Sentiment | radio | link |
2024 (9 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 1 | Semantic Textual Relatedness | likert, slider | link |
| Task 2 | Safe Biomedical NLI | radio, text | link |
| Task 3 | Multimodal Emotion Cause Analysis | span, radio | link |
| Task 4 | Persuasion Techniques in Memes | multiselect, span | link |
| Task 5 | Argument Reasoning in Civil Procedure | radio | link |
| Task 7 | NumEval - Numeral-Aware Language Understanding and Generation | radio, number, text | link |
| Task 8 | Machine-Generated Text Detection | radio | link |
| Task 9 | BRAINTEASER - Commonsense-Defying QA | radio, text | link |
| Task 10 | Emotion Discovery and Reasoning its Flip | radio, span | link |
2025 (10 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 1 | AdMIRe - Advancing Multimodal Idiomaticity Representation | radio | link |
| Task 2 | EA-MT - Entity-Aware Machine Translation | span, radio, text | link |
| Task 4 | Unlearning Sensitive Content from LLMs | radio, text, likert | link |
| Task 5 | LLMs4Subjects - Subject Tagging for a Technical Library | multiselect, radio | link |
| Task 6 | PromiseEval - Corporate ESG Promise Verification | radio, text | link |
| Task 7 | Multilingual and Crosslingual Fact-Checked Claim Retrieval | radio, text | link |
| Task 8 | QA over Tabular Data | text, radio | link |
| Task 9 | Food Hazard Detection | span, radio, multiselect | link |
| Task 10 | Multilingual Characterization and Extraction of Narratives | span, multiselect, text | link |
| Task 11 | Text-Based Emotion Detection | radio, multiselect | link |
2026 (10 tasks)
| Task | Title | Annotation Types | Directory |
|---|---|---|---|
| Task 2 | Emotional Valence and Arousal over Time from Ecological Essays | slider, text | link |
| Task 3 | DimABSA: Dimensional Aspect-Based Sentiment Analysis | span, slider, text | link |
| Task 4 | Narrative Story Similarity | radio, multiselect | link |
| Task 5 | Rating Plausibility of Word Senses through Narrative Understanding | likert | link |
| Task 6 | CLARITY: Unmasking Political Question Evasions | radio, text | link |
| Task 7 | Everyday Knowledge Across Diverse Languages and Cultures (BLEnD) | text, radio | link |
| Task 8 | MTRAGEval: Evaluating Multi-Turn RAG Conversations | radio, likert, text | link |
| Task 9 | POLAR: Detecting Multilingual, Multicultural and Multievent Online Polarization | radio, multiselect, text | link |
| Task 10 | PsyCoMark: Psycholinguistic Conspiracy Marker Extraction and Detection | radio, span | link |
| Task 13 | Detecting Machine-Generated Code | radio, text | link |
Task Count
Total: 100 SemEval task designs across 14 years