Parallel Processing Evaluated 1 Blueprint
Apache Spark logo

Apache Spark

Apache Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Java, Scala, Python, and R, and an optimized engine that supports general execution graphs. Spark can run in a variety of modes, including standalone, Mesos, YARN, or Kubernetes, and can process data from diverse sources like HDFS, Apache Cassandra, Apache HBase, and Amazon S3.

Category / Domain Parallel Processing
Reference Blueprints 1 Associated
I/O Performance Empirical Profile

Core Information

#core_information

Name, description, vendor website, logo, technology type and tags.

Overview & Role

Apache Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Java, Scala, Python, and R, and an optimized engine that supports general execution graphs. Spark can run in a variety of modes, including standalone, Mesos, YARN, or Kubernetes, and can process data from diverse sources like HDFS, Apache Cassandra, Apache HBase, and Amazon S3.

Domain Classification & Tags

Performance Profile

#performance_profile

Typical read and write throughput, payload size and processing latency.

Read Performance Profile

Throughput: Tens of GB/s to TB/s (depending on cluster size and data source)

Payload Size: MBs to GBs per partition/task

Processing Latency: Seconds to minutes (for batch jobs), milliseconds for streaming micro-batches

Write Performance Profile

Throughput: Tens of GB/s to TB/s (depending on cluster size and data sink)

Payload Size: MBs to GBs per partition/task

Processing Latency: Seconds to minutes (for batch jobs), milliseconds for streaming micro-batches

Architecture Diagram

#architecture_diagram

Reference diagram of the technology's internal architecture.

Component Architecture & Topology

Features

#features

Catalogued product capabilities and what each one does.

Empirical data for features is currently being compiled in the global catalog.

Typical Use Cases

#typical_use_cases

Scenarios the technology is commonly chosen for.

Empirical data for typical use cases is currently being compiled in the global catalog.

Known Customers

#known_customers

Publicly referenced organisations using the technology.

Empirical data for known customers is currently being compiled in the global catalog.

Known Integrations

#known_integrations

Other products and services it is documented to work with.

Empirical data for known integrations is currently being compiled in the global catalog.

Connectors

#connectors

Directional data connections to other technologies, with direction and maturity.

Empirical data for connectors is currently being compiled in the global catalog.

Reference Architectures

#reference_architectures

Published architectures where the technology is used or mentioned.

Security Features

#security_features

Built-in security and access-control capabilities.

Empirical data for security features is currently being compiled in the global catalog.

Known Issues

#known_issues

Documented limitations, defects and operational pitfalls.

Empirical data for known issues is currently being compiled in the global catalog.

Guidelines

#guidelines

Recommended practices for adopting and operating the technology.

Empirical data for guidelines is currently being compiled in the global catalog.

Standards & Compliance

#standards_and_compliance

Standards, certifications and control requirements it maps to.

Empirical data for standards & compliance is currently being compiled in the global catalog.

Sources

#sources

Documentation and research references behind the recorded information.

Empirical data for sources is currently being compiled in the global catalog.

Expert Validation

#expert_validation

Whether domain experts reviewed and confirmed the content.

Empirical data for expert validation is currently being compiled in the global catalog.