Apache Spark

E185661

Apache Spark is an open-source, distributed data processing engine designed for large-scale data analytics, machine learning, and stream processing.

All labels observed (12)

How this entity was disambiguated

Statements (91)

Predicate Object
instanceOf big data framework
cluster computing framework
distributed data processing engine
open-source software
abbreviation RDD
linked to: Apache Spark
architecture master-slave architecture
canRunOn Apache Mesos
Hadoop YARN
linked to: YARN

Kubernetes
standalone cluster manager
canUseStorage Amazon S3
Azure Data Lake Storage
Google Cloud Storage
Hadoop Distributed File System
linked to: HDFS

local file system
category big data analytics
data engineering
machine learning platform
stream processing framework
component GraphX
MLlib
linked to: Apache Spark

PySpark
Spark Core
Spark SQL
linked to: Apache Spark

Spark Streaming
linked to: Apache Spark

SparkR
linked to: Apache Spark

Structured Streaming
coreAbstraction Resilient Distributed Dataset
linked to: Apache Spark
designedFor batch processing
interactive data analytics
large-scale data processing
machine learning workloads
stream processing
developer Apache Software Foundation
donatedTo Apache Software Foundation
donationYear 2013
executionModel in-memory computing
hasComponent cluster manager
driver program
executors
initialReleaseDate 2010
integratesWith Apache Cassandra
Apache HBase
Apache Hadoop
linked to: Hadoop

Apache Hive
Apache Kafka
JDBC data sources
license Apache License 2.0
optimizedFor in-memory data processing
originatedAt UC Berkeley AMPLab
programmingLanguage Java
Python
R
SQL
Scala
provides Catalyst query optimizer
Tungsten execution engine
high-level APIs
low-level RDD API
schedulingUnit job
stage
task
supports SQL queries
batch processing
data parallelism
distributed computing
fault tolerance
graph processing
lazy evaluation
machine learning algorithms
stream processing
task parallelism
supportsAbstraction DataFrame
Dataset
supportsDeployment cloud environments
on-premises clusters
supportsLanguageAPI Java API
PySpark
Scala API
linked to: Scala

Spark SQL
linked to: Apache Spark

SparkR
linked to: Apache Spark
topLevelProjectSince 2014
useCase ETL pipelines
data warehousing
graph analytics
log processing
real-time analytics
recommendation systems
website https://spark.apache.org
writtenIn Java
Scala

How these facts were elicited

Referenced by (36)

Full triples — surface form annotated when it differs from this entity's canonical label.

Azure Synapse Analytics supports Apache Spark
Hadoop influenced Apache Spark
Scala ecosystem Apache Spark
KMeans implementedIn Apache Spark MLlib
linked to: Apache Spark
AWS Glue programmingModel Apache Spark
ORC usedIn Apache Spark
Avro usedWith Apache Spark
Apache Mesos supportsFramework Apache Spark
Apache Spark supportsLanguageAPI SparkR
linked to: Apache Spark
Apache Spark supportsLanguageAPI Spark SQL
linked to: Apache Spark
Apache Spark coreAbstraction Resilient Distributed Dataset
linked to: Apache Spark
Apache Spark abbreviation RDD
linked to: Apache Spark
Apache Spark component Spark SQL
linked to: Apache Spark
Apache Spark component Spark Streaming
linked to: Apache Spark
Apache Spark component MLlib
linked to: Apache Spark
Apache Spark component SparkR
linked to: Apache Spark
Synapse Studio supports Apache Spark
YARN supportsFramework Apache Spark
MapReduce influenced Apache Spark
Apache Storm competesWith Apache Spark Streaming
linked to: Apache Spark
Apache Hive runsOn Apache Spark
Apache HBase integratesWith Apache Spark
Apache Mahout integratesWith Apache Spark
Google MapReduce influenced Apache Spark programming model
linked to: Apache Spark
HDFS usedBy Apache Spark
Apache Pig executionEngine Spark
linked to: Apache Spark
Apache Pig comparedWith Apache Spark SQL
linked to: Apache Spark
NVIDIA RAPIDS integratesWith Apache Spark
Databricks coreTechnology Apache Spark
Apache Software Foundation governs Apache Spark
subject linked to: ASF
Apache Software Foundation hasKeyProject Apache Spark
subject linked to: ASF
ApacheCon isRelatedTo Apache Spark
Cloudera usesTechnology Apache Spark