Triple

T23506228
Position Surface form Disambiguated ID Type / Status
Subject Mirish languages E572288 entity
Predicate hasMember P10 FINISHED
Object Tangam language
Tangam language is a highly endangered Tibeto-Burman language spoken by a small indigenous community in Arunachal Pradesh, India.
E1589149 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Tangam language | Statement: [Mirish languages, hasMember, Tangam language]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Tangam language
Context triple: [Mirish languages, hasMember, Tangam language]
  • A. Tangsa language
    The Tangsa language is a Sino-Tibetan language spoken primarily by the Tangsa people in northeastern India and parts of Myanmar.
  • B. Xamtanga language
    Xamtanga is a Central Cushitic (Agaw) language spoken primarily in northern Ethiopia.
  • C. Ta-ang language
    The Ta-ang language is a Mon–Khmer language spoken primarily by the Ta-ang (Palaung) ethnic group in parts of Myanmar, China, and neighboring regions.
  • D. Tanema language
    Tanema is a nearly extinct Oceanic language once spoken on Vanikoro Island in the Temotu Province of the Solomon Islands.
  • E. Tumari language
    The Tumari language is a lesser-known Saharan language spoken by communities in parts of the central Sahara region of Africa.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Tangam language
Triple: [Mirish languages, hasMember, Tangam language]
Generated description
Tangam language is a highly endangered Tibeto-Burman language spoken by a small indigenous community in Arunachal Pradesh, India.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Tangam language
Target entity description: Tangam language is a highly endangered Tibeto-Burman language spoken by a small indigenous community in Arunachal Pradesh, India.
  • A. Tangsa language
    The Tangsa language is a Sino-Tibetan language spoken primarily by the Tangsa people in northeastern India and parts of Myanmar.
  • B. Xamtanga language
    Xamtanga is a Central Cushitic (Agaw) language spoken primarily in northern Ethiopia.
  • C. Ta-ang language
    The Ta-ang language is a Mon–Khmer language spoken primarily by the Ta-ang (Palaung) ethnic group in parts of Myanmar, China, and neighboring regions.
  • D. Tanema language
    Tanema is a nearly extinct Oceanic language once spoken on Vanikoro Island in the Temotu Province of the Solomon Islands.
  • E. Tumari language
    The Tumari language is a lesser-known Saharan language spoken by communities in parts of the central Sahara region of Africa.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69e245b5e4208190bac8a6509867e394 completed April 17, 2026, 2:37 p.m.
NER Named-entity recognition batch_69f1a9009ca081908c4cffb8c32293ec completed April 29, 2026, 6:45 a.m.
NED1 Entity disambiguation (via context triple) batch_6a0c8277b97481909b1405fd7469b910 completed May 19, 2026, 3:32 p.m.
NEDg Description generation batch_6a0ca6f2d4288190b0b5fd46d2a387bf completed May 19, 2026, 6:07 p.m.
NED2 Entity disambiguation (via description) batch_6a0ca7f09af48190b7bfdae25d2f286b completed May 19, 2026, 6:12 p.m.
Created at: April 17, 2026, 6:07 p.m.