Triple

T20899314
Position Surface form Disambiguated ID Type / Status
Subject Khakas language E514627 entity
Predicate hasDialects P4251 FINISHED
Object Kyzyl dialect
The Kyzyl dialect is a regional variety of the Khakas language spoken by Khakas communities in parts of Siberia, Russia.
E1457244 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Kyzyl dialect | Statement: [Khakas language, hasDialects, Kyzyl dialect]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Kyzyl dialect
Context triple: [Khakas language, hasDialects, Kyzyl dialect]
  • A. Saryk dialect
    The Saryk dialect is a regional variety of the Turkmen language traditionally spoken by the Saryk Turkmen people of Central Asia.
  • B. Buynaksk dialect
    The Buynaksk dialect is a regional variety of the Kumyk language spoken around the city of Buynaksk in Dagestan, Russia.
  • C. Khunzakh dialect
    The Khunzakh dialect is a regional variety of the Avar language spoken in and around the village of Khunzakh in Dagestan, Russia.
  • D. Khoshut dialect
    The Khoshut dialect is a regional variety of the Oirat Mongolic language traditionally spoken by the Khoshut people of Central Asia.
  • E. Oroqen dialect
    The Oroqen dialect is a regional variety of the Tungusic Evenki language spoken primarily by the Oroqen people of northeastern China.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Kyzyl dialect
Triple: [Khakas language, hasDialects, Kyzyl dialect]
Generated description
The Kyzyl dialect is a regional variety of the Khakas language spoken by Khakas communities in parts of Siberia, Russia.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Kyzyl dialect
Target entity description: The Kyzyl dialect is a regional variety of the Khakas language spoken by Khakas communities in parts of Siberia, Russia.
  • A. Saryk dialect
    The Saryk dialect is a regional variety of the Turkmen language traditionally spoken by the Saryk Turkmen people of Central Asia.
  • B. Buynaksk dialect
    The Buynaksk dialect is a regional variety of the Kumyk language spoken around the city of Buynaksk in Dagestan, Russia.
  • C. Khunzakh dialect
    The Khunzakh dialect is a regional variety of the Avar language spoken in and around the village of Khunzakh in Dagestan, Russia.
  • D. Khoshut dialect
    The Khoshut dialect is a regional variety of the Oirat Mongolic language traditionally spoken by the Khoshut people of Central Asia.
  • E. Oroqen dialect
    The Oroqen dialect is a regional variety of the Tungusic Evenki language spoken primarily by the Oroqen people of northeastern China.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69e0b4f8a1108190bce3d31331290ced completed April 16, 2026, 10:07 a.m.
NER Named-entity recognition batch_69e6e8f92bd88190b59b2131ad1d9aa1 completed April 21, 2026, 3:03 a.m.
NED1 Entity disambiguation (via context triple) batch_6a0918cefa2081909768f9923a96f209 completed May 17, 2026, 1:24 a.m.
NEDg Description generation batch_6a091ca90e788190bb74c161c9233634 completed May 17, 2026, 1:40 a.m.
NED2 Entity disambiguation (via description) batch_6a091d9402448190bb223e0950450e4e completed May 17, 2026, 1:44 a.m.
Created at: April 16, 2026, 12:47 p.m.