Triple
T36489678
| Position | Surface form | Disambiguated ID | Type / Status |
|---|---|---|---|
| Subject | WMT English-French dataset |
E899019
|
entity |
| Predicate | typicalPreprocessing |
P104012
|
FINISHED |
| Object | tokenization |
—
|
LITERAL FINISHED |
How this triple was built (2 steps)
Every LLM step that produced this triple, in pipeline order — named-entity classification, the disambiguation choices (the exact options shown, with the pick highlighted), and the generated description. The batch + timestamp of each is in the Provenance table below.
NER
Named-entity recognition
gpt-5-mini
Instruction
Given a phrase, classify it is english named entity (e.g., persons, organizations, works of art) in Latin script, or not (e.g., literals, dates, URLs, verbose phrases). For disambiguation, the statement where the phrase occurs as object is also given. Please return a JSON object with `phrase` (string, the phrase being analyzed) and `is_ne` (boolean, indicating whether the phrase is a Named Entity).
Input
Phrase: tokenization | Statement: [WMT English-French dataset, typicalPreprocessing, tokenization]
PD
Predicate disambiguation
gpt-5-mini-2025-08-07
Target predicate: typicalPreprocessing Context triple: [WMT English-French dataset, typicalPreprocessing, tokenization]
-
A.
typicalPreparation
Indicates the usual or standard way in which something is prepared or made.
-
B.
typicalProcess
chosen
Indicates that the related process is characteristic, usual, or commonly occurring for the given entity or context.
-
C.
typicalInitialization
Indicates the standard or commonly used way an entity is initially set up, configured, or brought into a valid starting state.
-
D.
processingUse
Indicates that one entity uses or applies another entity as part of a processing or transformation activity.
-
E.
standardBefore
Indicates that one standard must be satisfied, applied, or occur earlier in sequence or priority than another standard.
- F. None of above.
Provenance (3 batches)
The batch behind each pipeline step, in order, with when it ran. Timestamps are batch-level — stages were processed in waves, so the object chain (NER → NED1 → NEDg → NED2) reads in order, but predicate / elicitation batches can sit in a different wave.
| Step | Stage | Batch ID | Status | When |
|---|---|---|---|---|
| creating | Elicitation | batch_69f76e5ad4588190bdbce60c52fbb785 |
completed | May 3, 2026, 3:48 p.m. |
| NER | Named-entity recognition | batch_6a037c8e2c648190a65fc9c7872861af |
completed | May 12, 2026, 7:16 p.m. |
| PD | Predicate disambiguation | batch_6a037a0bf4b88190bdcfae9a14b51f0a |
completed | May 12, 2026, 7:05 p.m. |
Created at: May 3, 2026, 4:10 p.m.