Triple

T18784107
Position Surface form Disambiguated ID Type / Status
Subject modENCODE project E459329 entity
Predicate hasPublication P80 FINISHED
Object Science 2010 modENCODE integrative analysis papers
The Science 2010 modENCODE integrative analysis papers are a landmark set of publications that combined large-scale genomic, epigenomic, and transcriptomic data to comprehensively map functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
E459329 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Science 2010 modENCODE integrative analysis papers | Statement: [modENCODE project, hasPublication, Science 2010 modENCODE integrative analysis papers]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Science 2010 modENCODE integrative analysis papers
Context triple: [modENCODE project, hasPublication, Science 2010 modENCODE integrative analysis papers]
  • A. modENCODE project
    The modENCODE project is a large-scale genomics initiative aimed at comprehensively identifying and annotating functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
  • B. Genome Data Viewer
    Genome Data Viewer is an NCBI web-based genome browser that allows users to visualize, explore, and analyze annotated genomic sequences and features.
  • C. GEO DataSets
    GEO DataSets is a public NCBI repository that organizes and provides access to curated gene expression and other functional genomics experiment data.
  • D. The Language of the Genes
    The Language of the Genes is a popular science book by geneticist Steve Jones that explains how genetics shapes human evolution, diversity, and behavior in accessible terms.
  • E. Edico Genome
    Edico Genome was a biotechnology company specializing in high-speed genomic data analysis through its DRAGEN bio-IT platform and FPGA-based acceleration technology.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Science 2010 modENCODE integrative analysis papers
Triple: [modENCODE project, hasPublication, Science 2010 modENCODE integrative analysis papers]
Generated description
The Science 2010 modENCODE integrative analysis papers are a landmark set of publications that combined large-scale genomic, epigenomic, and transcriptomic data to comprehensively map functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Science 2010 modENCODE integrative analysis papers
Target entity description: The Science 2010 modENCODE integrative analysis papers are a landmark set of publications that combined large-scale genomic, epigenomic, and transcriptomic data to comprehensively map functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
  • A. modENCODE project chosen
    The modENCODE project is a large-scale genomics initiative aimed at comprehensively identifying and annotating functional elements in model organisms such as Drosophila melanogaster and Caenorhabditis elegans.
  • B. Genome Data Viewer
    Genome Data Viewer is an NCBI web-based genome browser that allows users to visualize, explore, and analyze annotated genomic sequences and features.
  • C. GEO DataSets
    GEO DataSets is a public NCBI repository that organizes and provides access to curated gene expression and other functional genomics experiment data.
  • D. The Language of the Genes
    The Language of the Genes is a popular science book by geneticist Steve Jones that explains how genetics shapes human evolution, diversity, and behavior in accessible terms.
  • E. Edico Genome
    Edico Genome was a biotechnology company specializing in high-speed genomic data analysis through its DRAGEN bio-IT platform and FPGA-based acceleration technology.
  • F. None of above.

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d8d396f54c8190ba49db31e8743842 completed April 10, 2026, 10:40 a.m.
NER Named-entity recognition batch_69e5977ffa648190be5f682bba47fb0a completed April 20, 2026, 3:03 a.m.
NED1 Entity disambiguation (via context triple) batch_6a05471ae8fc8190a18870af6b487105 completed May 14, 2026, 3:52 a.m.
NEDg Description generation batch_6a054abe671c8190b9ef2d324fe23a96 completed May 14, 2026, 4:08 a.m.
NED2 Entity disambiguation (via description) batch_6a054b2bf4848190ae3329f47c30210c completed May 14, 2026, 4:10 a.m.
Created at: April 10, 2026, 11:52 a.m.