Apache Spark
Apache Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Java, Scala, Python, and R, and an optimized engine that supports general execution graphs. Spark can run in a variety of modes, including standalone, Mesos, YARN, or Kubernetes, and can process data from diverse sources like HDFS, Apache Cassandra, Apache HBase, and Amazon S3.
Core Information
#core_informationName, description, vendor website, logo, technology type and tags.
Overview & Role
Apache Spark is a unified analytics engine for large-scale data processing. It provides high-level APIs in Java, Scala, Python, and R, and an optimized engine that supports general execution graphs. Spark can run in a variety of modes, including standalone, Mesos, YARN, or Kubernetes, and can process data from diverse sources like HDFS, Apache Cassandra, Apache HBase, and Amazon S3.
Domain Classification & Tags
Performance Profile
#performance_profileTypical read and write throughput, payload size and processing latency.
Read Performance Profile
Throughput: Tens of GB/s to TB/s (depending on cluster size and data source)
Payload Size: MBs to GBs per partition/task
Processing Latency: Seconds to minutes (for batch jobs), milliseconds for streaming micro-batches
Write Performance Profile
Throughput: Tens of GB/s to TB/s (depending on cluster size and data sink)
Payload Size: MBs to GBs per partition/task
Processing Latency: Seconds to minutes (for batch jobs), milliseconds for streaming micro-batches
Architecture Diagram
#architecture_diagramReference diagram of the technology's internal architecture.
Component Architecture & Topology
Features
#featuresCatalogued product capabilities and what each one does.
Empirical data for features is currently being compiled in the global catalog.
Typical Use Cases
#typical_use_casesScenarios the technology is commonly chosen for.
Empirical data for typical use cases is currently being compiled in the global catalog.
Known Customers
#known_customersPublicly referenced organisations using the technology.
Empirical data for known customers is currently being compiled in the global catalog.
Known Integrations
#known_integrationsOther products and services it is documented to work with.
Empirical data for known integrations is currently being compiled in the global catalog.
Connectors
#connectorsDirectional data connections to other technologies, with direction and maturity.
Empirical data for connectors is currently being compiled in the global catalog.
Reference Architectures
#reference_architecturesPublished architectures where the technology is used or mentioned.
Associated with 1 verified reference architecture blueprint:
Security Features
#security_featuresBuilt-in security and access-control capabilities.
Empirical data for security features is currently being compiled in the global catalog.
Known Issues
#known_issuesDocumented limitations, defects and operational pitfalls.
Empirical data for known issues is currently being compiled in the global catalog.
Guidelines
#guidelinesRecommended practices for adopting and operating the technology.
Empirical data for guidelines is currently being compiled in the global catalog.
Standards & Compliance
#standards_and_complianceStandards, certifications and control requirements it maps to.
Empirical data for standards & compliance is currently being compiled in the global catalog.
Sources
#sourcesDocumentation and research references behind the recorded information.
Empirical data for sources is currently being compiled in the global catalog.
Expert Validation
#expert_validationWhether domain experts reviewed and confirmed the content.
Empirical data for expert validation is currently being compiled in the global catalog.