Triple

T19471892
Position Surface form Disambiguated ID Type / Status
Subject Kuba dialect E487141 entity
Predicate hasAncestor P369 FINISHED
Object Proto-Lezgic
Proto-Lezgic is the reconstructed common ancestor of the Lezgic branch of Northeast Caucasian languages, from which varieties such as the Kuba dialect ultimately developed.
E1378246 NE FINISHED

How this triple was built (4 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: Proto-Lezgic | Statement: [Kuba dialect, hasAncestor, Proto-Lezgic]
NED1 Entity disambiguation (via context triple) gpt-5-mini-2025-08-07
Target entity: Proto-Lezgic
Context triple: [Kuba dialect, hasAncestor, Proto-Lezgic]
  • A. Proto-Gur
    Proto-Gur is the reconstructed common ancestor of the Gur languages of West Africa, inferred through comparative linguistic analysis.
  • B. Proto-Dargwa
    Proto-Dargwa is the reconstructed ancestral language from which the modern Dargwa varieties of the Northeast Caucasian language family are derived.
  • C. Proto-Koman
    Proto-Koman is the reconstructed ancestral language from which the modern Koman languages of the Nilo-Saharan family are derived.
  • D. Proto-Daju language
    Proto-Daju language is the reconstructed common ancestor of the Daju languages, hypothesized through comparative linguistic analysis.
  • E. Proto-Mari
    Proto-Mari is the reconstructed ancestral language from which the modern Hill Mari language developed.
  • F. None of above. chosen
  • G. Unsure - the case is ambiguous/there is not enough information to decide.
NEDg Description generation gpt-5.1
Instruction
Generate a one-sentence description of the target entity. 
You are given a context triple in the form (subject, predicate, object), where the object is the target entity. 
# Instructions
Use the triple to infer relevant information about the entity. Describe the entity based on what is most defining, well-known. 
Avoid repeating the information from the triple, unless really essential.
# Response Format
Return only the sentence: "Description: [one-sentence description of the target entity]"
Input
Entity: Proto-Lezgic
Triple: [Kuba dialect, hasAncestor, Proto-Lezgic]
Generated description
Proto-Lezgic is the reconstructed common ancestor of the Lezgic branch of Northeast Caucasian languages, from which varieties such as the Kuba dialect ultimately developed.
NED2 Entity disambiguation (via description) gpt-5-mini-2025-08-07
Target entity: Proto-Lezgic
Target entity description: Proto-Lezgic is the reconstructed common ancestor of the Lezgic branch of Northeast Caucasian languages, from which varieties such as the Kuba dialect ultimately developed.
  • A. Proto-Gur
    Proto-Gur is the reconstructed common ancestor of the Gur languages of West Africa, inferred through comparative linguistic analysis.
  • B. Proto-Dargwa
    Proto-Dargwa is the reconstructed ancestral language from which the modern Dargwa varieties of the Northeast Caucasian language family are derived.
  • C. Proto-Koman
    Proto-Koman is the reconstructed ancestral language from which the modern Koman languages of the Nilo-Saharan family are derived.
  • D. Proto-Daju language
    Proto-Daju language is the reconstructed common ancestor of the Daju languages, hypothesized through comparative linguistic analysis.
  • E. Proto-Mari
    Proto-Mari is the reconstructed ancestral language from which the modern Hill Mari language developed.
  • F. None of above. chosen

Provenance (5 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69d8e8d924388190b847cb15bb3d0aff completed April 10, 2026, 12:11 p.m.
NER Named-entity recognition batch_69e633e8d9188190a4939f03bad89add completed April 20, 2026, 2:10 p.m.
NED1 Entity disambiguation (via context triple) batch_6a07404fd4348190bc2efbebc7828716 completed May 15, 2026, 3:48 p.m.
NEDg Description generation batch_6a074155addc8190a80e37d8f9d537a8 completed May 15, 2026, 3:52 p.m.
NED2 Entity disambiguation (via description) batch_6a0741e454a48190ac03aa458ca2b3b3 completed May 15, 2026, 3:55 p.m.
Created at: April 10, 2026, 1:39 p.m.