Hadoop

E35621

Hadoop is an open-source framework that enables distributed storage and parallel processing of large data sets across clusters of commodity hardware.

All labels observed (11)

How this entity was disambiguated

Statements (60)

Predicate Object
instanceOf big data framework
distributed computing framework
open-source software framework
developer Apache Software Foundation
domain big data
ecosystemIncludes Apache Flume
Apache HBase
Apache Hive
Apache Mahout
Apache Oozie
Apache Pig
Apache Sqoop
Apache ZooKeeper
hasComponent HDFS
Hadoop Common
linked to: Hadoop

Hadoop Distributed File System
linked to: Hadoop

MapReduce
YARN
Yet Another Resource Negotiator
influenced Apache Flink
Apache Spark
Apache Storm
initiallyInspiredBy Google File System
Google MapReduce
license Apache License 2.0
operatingSystem Cross-platform
partOf Apache Hadoop ecosystem
processingLayer MapReduce
programmingLanguage Java
resourceManagementLayer YARN
runsOn clusters of commodity hardware
storageLayer HDFS
supportsArchitecture master-slave architecture
supportsDataReplicationFactor configurable replication factor
supportsFeature batch processing
data replication
distributed storage
fault tolerance
horizontal scalability
parallel processing
supportsHighAvailability NameNode high availability
supportsModel MapReduce programming model
supportsOperatingSystem Linux
Windows
macOS
supportsProgrammingLanguage C++
Java
Python
R
Scala
supportsSecurity Kerberos-based authentication
useCase ETL workloads
data warehousing
large-scale data processing
log processing
machine learning at scale
writtenIn C
Java
Python
Shell
linked to: Unix shell

How these facts were elicited

Referenced by (36)

Full triples — surface form annotated when it differs from this entity's canonical label.

Hadoop hasComponent Hadoop Common
linked to: Hadoop
Hadoop hasComponent Hadoop Distributed File System
linked to: Hadoop
Apache Software Foundation overseesProject Apache Hadoop
linked to: Hadoop
ORC usedIn Hadoop ecosystem
linked to: Hadoop
Google Cloud Dataproc supportsFramework Apache Hadoop
linked to: Hadoop
Avro partOf Apache Hadoop ecosystem
linked to: Hadoop
Avro usedWith Apache Hadoop
linked to: Hadoop
Apache Mesos supportsFramework Apache Hadoop
linked to: Hadoop
Apache Spark integratesWith Apache Hadoop
linked to: Hadoop
Yet Another Resource Negotiator partOf Apache Hadoop ecosystem
linked to: Hadoop
Yet Another Resource Negotiator introducedIn Hadoop 2.x
linked to: Hadoop
Yet Another Resource Negotiator developedAsPartOf Apache Hadoop project
linked to: Hadoop
YARN partOf Apache Hadoop ecosystem
linked to: Hadoop
YARN introducedIn Apache Hadoop 2.x
linked to: Hadoop
MapReduce influenced Apache Hadoop MapReduce
linked to: Hadoop
Apache Storm supportsIntegrationWith Apache Hadoop
linked to: Hadoop
Apache Hive partOf Apache Hadoop ecosystem
linked to: Hadoop
Apache HBase runsOnTopOf Apache Hadoop
linked to: Hadoop
Apache Oozie integratesWith Apache Hadoop
linked to: Hadoop
Apache Oozie supportsVersion Hadoop 1.x
linked to: Hadoop
Apache Oozie supportsVersion Hadoop 2.x
linked to: Hadoop
Apache ZooKeeper usedBy Apache Hadoop
linked to: Hadoop
Apache Sqoop supportsPlatform Apache Hadoop
linked to: Hadoop
Apache Flume supports Hadoop
Apache Mahout integratesWith Apache Hadoop
linked to: Hadoop
Google MapReduce influenced Apache Hadoop MapReduce
linked to: Hadoop
Apache Flink integratesWith Apache Hadoop
linked to: Hadoop
Apache Pig runsOn Hadoop
Apache Pig ecosystem Hadoop ecosystem
linked to: Hadoop
Apache Software Foundation governs Apache Hadoop
subject linked to: ASF
linked to: Hadoop
Apache Software Foundation hasKeyProject Apache Hadoop
subject linked to: ASF
linked to: Hadoop
ApacheCon isRelatedTo Apache Hadoop
linked to: Hadoop
Cloudera usesTechnology Apache Hadoop
linked to: Hadoop
Cloudera basedOn Apache Hadoop ecosystem
linked to: Hadoop