Triple

T36489678
Position Surface form Disambiguated ID Type / Status
Subject WMT English-French dataset E899019 entity
Predicate typicalPreprocessing P104012 FINISHED
Object tokenization LITERAL FINISHED

How this triple was built (2 steps)

Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.

NER Named-entity recognition gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: tokenization | Statement: [WMT English-French dataset, typicalPreprocessing, tokenization]
PD Predicate disambiguation gpt-5-mini-2025-08-07
Target predicate: typicalPreprocessing
Context triple: [WMT English-French dataset, typicalPreprocessing, tokenization]
  • A. typicalPreparation
    Indicates the usual or standard way in which something is prepared or made.
  • B. typicalProcess chosen
    Indicates that the related process is characteristic, usual, or commonly occurring for the given entity or context.
  • C. typicalInitialization
    Indicates the standard or commonly used way an entity is initially set up, configured, or brought into a valid starting state.
  • D. processingUse
    Indicates that one entity uses or applies another entity as part of a processing or transformation activity.
  • E. standardBefore
    Indicates that one standard must be satisfied, applied, or occur earlier in sequence or priority than another standard.
  • F. None of above.

Provenance (3 batches)

The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.

Step Stage Batch ID Status When
creating Elicitation batch_69f76e5ad4588190bdbce60c52fbb785 completed May 3, 2026, 3:48 p.m.
NER Named-entity recognition batch_6a037c8e2c648190a65fc9c7872861af completed May 12, 2026, 7:16 p.m.
PD Predicate disambiguation batch_6a037a0bf4b88190bdcfae9a14b51f0a completed May 12, 2026, 7:05 p.m.
Created at: May 3, 2026, 4:10 p.m.