Apache Spark

E185661

Apache Spark is an open-source, distributed data processing engine designed for large-scale data analytics, machine learning, and stream processing.

All labels observed (22)

How this entity was disambiguated

Statements (91)

Predicate Object
instanceOf big data framework ⓘ
cluster computing framework ⓘ
distributed data processing engine ⓘ
open-source software ⓘ
abbreviation RDD ⓘ
linked to: Apache Spark
architecture master-slave architecture ⓘ
canRunOn Apache Mesos ⓘ
Hadoop YARN ⓘ
linked to: YARN

Kubernetes ⓘ
standalone cluster manager ⓘ
canUseStorage Amazon S3 ⓘ
Azure Data Lake Storage ⓘ
Google Cloud Storage ⓘ
Hadoop Distributed File System ⓘ
linked to: HDFS

local file system ⓘ
category big data analytics ⓘ
data engineering ⓘ
machine learning platform ⓘ
stream processing framework ⓘ
component GraphX ⓘ
MLlib ⓘ
linked to: Apache Spark

PySpark ⓘ
Spark Core ⓘ
Spark SQL ⓘ
linked to: Apache Spark

Spark Streaming ⓘ
linked to: Apache Spark

SparkR ⓘ
linked to: Apache Spark

Structured Streaming ⓘ
coreAbstraction Resilient Distributed Dataset ⓘ
linked to: Apache Spark
designedFor batch processing ⓘ
interactive data analytics ⓘ
large-scale data processing ⓘ
machine learning workloads ⓘ
stream processing ⓘ
developer Apache Software Foundation ⓘ
donatedTo Apache Software Foundation ⓘ
donationYear 2013 ⓘ
executionModel in-memory computing ⓘ
hasComponent cluster manager ⓘ
driver program ⓘ
executors ⓘ
initialReleaseDate 2010 ⓘ
integratesWith Apache Cassandra ⓘ
Apache HBase ⓘ
Apache Hadoop ⓘ
linked to: Hadoop

Apache Hive ⓘ
Apache Kafka ⓘ
JDBC data sources ⓘ
license Apache License 2.0 ⓘ
optimizedFor in-memory data processing ⓘ
originatedAt UC Berkeley AMPLab ⓘ
programmingLanguage Java ⓘ
Python ⓘ
R ⓘ
SQL ⓘ
Scala ⓘ
provides Catalyst query optimizer ⓘ
Tungsten execution engine ⓘ
high-level APIs ⓘ
low-level RDD API ⓘ
schedulingUnit job ⓘ
stage ⓘ
task ⓘ
supports SQL queries ⓘ
batch processing ⓘ
data parallelism ⓘ
distributed computing ⓘ
fault tolerance ⓘ
graph processing ⓘ
lazy evaluation ⓘ
machine learning algorithms ⓘ
stream processing ⓘ
task parallelism ⓘ
supportsAbstraction DataFrame ⓘ
Dataset ⓘ
supportsDeployment cloud environments ⓘ
on-premises clusters ⓘ
supportsLanguageAPI Java API ⓘ
PySpark ⓘ
Scala API ⓘ
linked to: Scala

Spark SQL ⓘ
linked to: Apache Spark

SparkR ⓘ
linked to: Apache Spark
topLevelProjectSince 2014 ⓘ
useCase ETL pipelines ⓘ
data warehousing ⓘ
graph analytics ⓘ
log processing ⓘ
real-time analytics ⓘ
recommendation systems ⓘ
website https://spark.apache.org ⓘ
writtenIn Java ⓘ
Scala ⓘ

How these facts were elicited

Referenced by (85)

Full triples — surface form annotated when it differs from this entity's canonical label.

Azure Synapse Analytics → supports → Apache Spark ⓘ
Hadoop → influenced → Apache Spark ⓘ
Scala → ecosystem → Apache Spark ⓘ
KMeans → implementedIn → Apache Spark MLlib ⓘ
linked to: Apache Spark
AWS Glue → programmingModel → Apache Spark ⓘ
ORC → usedIn → Apache Spark ⓘ
Avro → usedWith → Apache Spark ⓘ
Apache Mesos → supportsFramework → Apache Spark ⓘ
Apache Spark → supportsLanguageAPI → SparkR ⓘ
linked to: Apache Spark
Apache Spark → supportsLanguageAPI → Spark SQL ⓘ
linked to: Apache Spark
Apache Spark → coreAbstraction → Resilient Distributed Dataset ⓘ
linked to: Apache Spark
Apache Spark → abbreviation → RDD ⓘ
linked to: Apache Spark
Apache Spark → component → Spark SQL ⓘ
linked to: Apache Spark
Apache Spark → component → Spark Streaming ⓘ
linked to: Apache Spark
Apache Spark → component → MLlib ⓘ
linked to: Apache Spark
Apache Spark → component → SparkR ⓘ
linked to: Apache Spark
Synapse Studio → supports → Apache Spark ⓘ
YARN → supportsFramework → Apache Spark ⓘ
MapReduce → influenced → Apache Spark ⓘ
Apache Storm → competesWith → Apache Spark Streaming ⓘ
linked to: Apache Spark
Apache Hive → runsOn → Apache Spark ⓘ
Apache HBase → integratesWith → Apache Spark ⓘ
Apache Mahout → integratesWith → Apache Spark ⓘ
Google MapReduce → influenced → Apache Spark programming model ⓘ
linked to: Apache Spark
HDFS → usedBy → Apache Spark ⓘ
Apache Pig → executionEngine → Spark ⓘ
linked to: Apache Spark
Apache Pig → comparedWith → Apache Spark SQL ⓘ
linked to: Apache Spark
NVIDIA RAPIDS → integratesWith → Apache Spark ⓘ
Databricks → coreTechnology → Apache Spark ⓘ
Apache Software Foundation → governs → Apache Spark ⓘ
subject linked to: ASF
Apache Software Foundation → hasKeyProject → Apache Spark ⓘ
subject linked to: ASF
ApacheCon → isRelatedTo → Apache Spark ⓘ
Cloudera → usesTechnology → Apache Spark ⓘ
Apache ORC → usedIn → Apache Spark ⓘ
subject linked to: Apache ORC project
Apache Beam → supportsRunner → Apache Spark ⓘ
XGBoost → supportsLanguageBinding → Spark ⓘ
linked to: Apache Spark
Ray → integratesWith → Apache Spark ⓘ
Apache Airflow → integratesWith → Apache Spark ⓘ
Apache Parquet → usedWith → Apache Spark ⓘ
Amazon EMR → supportsFramework → Apache Spark ⓘ
Oracle Data Flow → supportsFramework → Apache Spark ⓘ
Oracle Big Data Service → supports → Apache Spark ⓘ
IBM Streams → alternativeTo → Apache Spark Streaming ⓘ
linked to: Apache Spark
Sahara API → supportsTechnology → Apache Spark ⓘ
OpenStack Sahara → supportsFramework → Apache Spark ⓘ
Apache Mesos → supportsFramework → Apache Spark ⓘ
subject linked to: Mesos