Hadoop

E35621

Hadoop is an open-source framework that enables distributed storage and parallel processing of large data sets across clusters of commodity hardware.

AI illustration

How this image was made

AI-generated illustration of Hadoop

This AI-generated illustration was produced by black-forest-labs/FLUX.2-dev (1024x1024) from a prompt written by openai/gpt-oss-120b from the entity's label + description.

Prompt

Generate an image of Hadoop (Hadoop is an open-source framework that enables distributed storage and parallel processing of large data sets across clusters of commodity hardware.)

All labels observed (17)

How this entity was disambiguated

Statements (60)

Predicate Object
instanceOf big data framework ⓘ
distributed computing framework ⓘ
open-source software framework ⓘ
developer Apache Software Foundation ⓘ
domain big data ⓘ
ecosystemIncludes Apache Flume ⓘ
Apache HBase ⓘ
Apache Hive ⓘ
Apache Mahout ⓘ
Apache Oozie ⓘ
Apache Pig ⓘ
Apache Sqoop ⓘ
Apache ZooKeeper ⓘ
hasComponent HDFS ⓘ
Hadoop Common ⓘ
linked to: Hadoop

Hadoop Distributed File System ⓘ
linked to: Hadoop

MapReduce ⓘ
YARN ⓘ
Yet Another Resource Negotiator ⓘ
influenced Apache Flink ⓘ
Apache Spark ⓘ
Apache Storm ⓘ
initiallyInspiredBy Google File System ⓘ
Google MapReduce ⓘ
license Apache License 2.0 ⓘ
operatingSystem Cross-platform ⓘ
partOf Apache Hadoop ecosystem ⓘ
processingLayer MapReduce ⓘ
programmingLanguage Java ⓘ
resourceManagementLayer YARN ⓘ
runsOn clusters of commodity hardware ⓘ
storageLayer HDFS ⓘ
supportsArchitecture master-slave architecture ⓘ
supportsDataReplicationFactor configurable replication factor ⓘ
supportsFeature batch processing ⓘ
data replication ⓘ
distributed storage ⓘ
fault tolerance ⓘ
horizontal scalability ⓘ
parallel processing ⓘ
supportsHighAvailability NameNode high availability ⓘ
supportsModel MapReduce programming model ⓘ
supportsOperatingSystem Linux ⓘ
Windows ⓘ
macOS ⓘ
supportsProgrammingLanguage C++ ⓘ
Java ⓘ
Python ⓘ
R ⓘ
Scala ⓘ
supportsSecurity Kerberos-based authentication ⓘ
useCase ETL workloads ⓘ
data warehousing ⓘ
large-scale data processing ⓘ
log processing ⓘ
machine learning at scale ⓘ
writtenIn C ⓘ
Java ⓘ
Python ⓘ
Shell ⓘ
linked to: Unix shell

How these facts were elicited

Referenced by (81)

Full triples — surface form annotated when it differs from this entity's canonical label.

Hadoop → hasComponent → Hadoop Common ⓘ
linked to: Hadoop
Hadoop → hasComponent → Hadoop Distributed File System ⓘ
linked to: Hadoop
Apache Software Foundation → overseesProject → Apache Hadoop ⓘ
linked to: Hadoop
ORC → usedIn → Hadoop ecosystem ⓘ
linked to: Hadoop
Google Cloud Dataproc → supportsFramework → Apache Hadoop ⓘ
linked to: Hadoop
Avro → partOf → Apache Hadoop ecosystem ⓘ
linked to: Hadoop
Avro → usedWith → Apache Hadoop ⓘ
linked to: Hadoop
Apache Mesos → supportsFramework → Apache Hadoop ⓘ
linked to: Hadoop
Apache Spark → integratesWith → Apache Hadoop ⓘ
linked to: Hadoop
Yet Another Resource Negotiator → partOf → Apache Hadoop ecosystem ⓘ
linked to: Hadoop
Yet Another Resource Negotiator → introducedIn → Hadoop 2.x ⓘ
linked to: Hadoop
Yet Another Resource Negotiator → developedAsPartOf → Apache Hadoop project ⓘ
linked to: Hadoop
YARN → partOf → Apache Hadoop ecosystem ⓘ
linked to: Hadoop
YARN → introducedIn → Apache Hadoop 2.x ⓘ
linked to: Hadoop
MapReduce → influenced → Apache Hadoop MapReduce ⓘ
linked to: Hadoop
Apache Storm → supportsIntegrationWith → Apache Hadoop ⓘ
linked to: Hadoop
Apache Hive → partOf → Apache Hadoop ecosystem ⓘ
linked to: Hadoop
Apache HBase → runsOnTopOf → Apache Hadoop ⓘ
linked to: Hadoop
Apache Oozie → integratesWith → Apache Hadoop ⓘ
linked to: Hadoop
Apache Oozie → supportsVersion → Hadoop 1.x ⓘ
linked to: Hadoop
Apache Oozie → supportsVersion → Hadoop 2.x ⓘ
linked to: Hadoop
Apache ZooKeeper → usedBy → Apache Hadoop ⓘ
linked to: Hadoop
Apache Sqoop → supportsPlatform → Apache Hadoop ⓘ
linked to: Hadoop
Apache Flume → supports → Hadoop ⓘ
Apache Mahout → integratesWith → Apache Hadoop ⓘ
linked to: Hadoop
Google MapReduce → influenced → Apache Hadoop MapReduce ⓘ
linked to: Hadoop
Apache Flink → integratesWith → Apache Hadoop ⓘ
linked to: Hadoop
Apache Pig → runsOn → Hadoop ⓘ
Apache Pig → ecosystem → Hadoop ecosystem ⓘ
linked to: Hadoop
Apache Software Foundation → governs → Apache Hadoop ⓘ
subject linked to: ASF
linked to: Hadoop
Apache Software Foundation → hasKeyProject → Apache Hadoop ⓘ
subject linked to: ASF
linked to: Hadoop
ApacheCon → isRelatedTo → Apache Hadoop ⓘ
linked to: Hadoop
Cloudera → usesTechnology → Apache Hadoop ⓘ
linked to: Hadoop
Cloudera → basedOn → Apache Hadoop ecosystem ⓘ
linked to: Hadoop
Optimized Row Columnar → usedIn → Apache Hadoop ecosystem ⓘ
linked to: Hadoop
Apache ORC → usedIn → Apache Hadoop ecosystem ⓘ
subject linked to: Apache ORC project
linked to: Hadoop
Facebook Engineering → openSourceContributorTo → Apache Hadoop ⓘ
linked to: Hadoop
Apache Airflow → integratesWith → Apache Hadoop ⓘ
linked to: Hadoop
Apache Parquet → usedWith → Apache Hadoop ⓘ
linked to: Hadoop
Apache Parquet → designedFor → Hadoop ecosystem ⓘ
linked to: Hadoop
Amazon EMR → supportsFramework → Apache Hadoop ⓘ
linked to: Hadoop
Oracle Big Data Service → supports → Apache Hadoop ⓘ
linked to: Hadoop
Oracle Big Data Appliance → uses → Apache Hadoop ⓘ
linked to: Hadoop
Sahara → supportsTechnology → Hadoop ⓘ
Sahara → supportsTechnology → Vanilla Hadoop ⓘ
linked to: Hadoop
OpenStack Sahara → supportsFramework → Apache Hadoop ⓘ
linked to: Hadoop