WARC

E457874

WARC is a standardized file format used to store and archive web crawls and their associated metadata at scale.

All labels observed (1)

Label Occurrences
WARC canonical 1

How this entity was disambiguated

Statements (48)

Predicate Object
instanceOf file format ⓘ
web archiving format ⓘ
allows linking between records via identifiers ⓘ
storing multiple representations of same resource ⓘ
associatedSoftware Heritrix web crawler ⓘ
OpenWayback ⓘ
Webrecorder tools ⓘ
pywb ⓘ
compatibleWith HTTP ⓘ
compressionSupport GZIP ⓘ
linked to: gzip
designedBy International Internet Preservation Consortium ⓘ
designedFor batch-oriented processing ⓘ
large-scale web crawls ⓘ
domain digital preservation ⓘ
web archiving ⓘ
fileExtension .warc ⓘ
.warc.gz ⓘ
fullName Web ARChive format ⓘ
governedBy ISO 28500:2009 ⓘ
ISO 28500:2017 ⓘ
linked to: ISO 28500
headerFormat text-based key-value headers ⓘ
initialPublicationYear 2009 ⓘ
latestRevisionYear 2017 ⓘ
mediaType application/warc ⓘ
payloadFormat binary or text payloads ⓘ
predecessor ARC file format ⓘ
primaryUse preserving web content at scale ⓘ
storing web crawls ⓘ
web archiving ⓘ
recordIdentification URI-based identifiers ⓘ
recordStructure sequence of self-contained records ⓘ
standardizedBy International Organization for Standardization ⓘ
standardNumber ISO 28500 ⓘ
status international standard ⓘ
stores HTTP request records ⓘ
HTTP response records ⓘ
continuation records ⓘ
conversion records ⓘ
metadata records ⓘ
resource records ⓘ
revisit records ⓘ
supports deduplication via revisit records ⓘ
embedded metadata ⓘ
long-term preservation of web content ⓘ
usedBy Internet Archive ⓘ
national libraries ⓘ
research institutions ⓘ
web archiving projects ⓘ

How these facts were elicited

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

Common Crawl → dataFormat → WARC ⓘ