Apache Gobblin

E705297

Apache Gobblin is an open-source distributed data integration framework designed for large-scale data ingestion, replication, and lifecycle management across diverse data sources and sinks.

All labels observed (1)

Label Occurrences
Apache Gobblin canonical 1

How this entity was disambiguated

Statements (53)

Predicate Object
instanceOf Apache Software Foundation project ⓘ
distributed data ingestion framework ⓘ
open-source data integration framework ⓘ
developer Apache Software Foundation ⓘ
donatedTo Apache Software Foundation ⓘ
feature checkpointing ⓘ
config-driven job specification ⓘ
fault tolerance ⓘ
job scheduling ⓘ
metrics collection ⓘ
monitoring and alerting ⓘ
pluggable source and sink architecture ⓘ
schema management ⓘ
task parallelism ⓘ
watermarking ⓘ
license Apache License 2.0 ⓘ
originatedAt LinkedIn ⓘ
programmingLanguage Java ⓘ
repository https://github.com/apache/gobblin ⓘ
supportsDataSourceType NoSQL databases ⓘ
RDBMS ⓘ
REST APIs ⓘ
file systems ⓘ
message queues ⓘ
supportsDeploymentModel MapReduce-based deployment ⓘ
YARN-based deployment ⓘ
cluster mode ⓘ
containerized deployment ⓘ
service mode ⓘ
standalone mode ⓘ
supportsEnvironment cloud deployments ⓘ
hybrid deployments ⓘ
on-premises deployments ⓘ
supportsPlatform Apache Helix ⓘ
Apache Kafka ⓘ
Apache YARN ⓘ
linked to: YARN

Hadoop ⓘ
Kubernetes ⓘ
supportsSinkType HDFS ⓘ
Kafka ⓘ
RDBMS ⓘ
data warehouses ⓘ
object stores ⓘ
supportsUseCase ETL ⓘ
data compaction ⓘ
data integration ⓘ
data lifecycle management ⓘ
data migration ⓘ
data quality management ⓘ
data replication ⓘ
large-scale data ingestion ⓘ
metadata management ⓘ
website https://gobblin.apache.org/ ⓘ

How these facts were elicited

Referenced by (1)

Full triples — surface form annotated when it differs from this entity's canonical label.

Apache Sqoop → supersededBy → Apache Gobblin ⓘ