Delta Lake
Delta Lake is an open-source storage layer that brings ACID transactions to Apache Spark and big data workloads. It enables building a Lakehouse architecture on top of existing data lakes, providing reliability, security, and performance for both streaming and batch operations. Key features include ACID transactions, scalable metadata handling, and unified streaming and batch data processing.
Core Information
#core_informationName, description, vendor website, logo, technology type and tags.
Overview & Role
Delta Lake is an open-source storage layer that brings ACID transactions to Apache Spark and big data workloads. It enables building a Lakehouse architecture on top of existing data lakes, providing reliability, security, and performance for both streaming and batch operations. Key features include ACID transactions, scalable metadata handling, and unified streaming and batch data processing.
Domain Classification & Tags
Performance Profile
#performance_profileTypical read and write throughput, payload size and processing latency.
Read Performance Profile
Throughput: Varies greatly with data size and query complexity, potentially hundreds of GB/s for large scans on optimized clusters
Payload Size: Varies from single rows (KB) to full tables (TB)
Processing Latency: Sub-second for small queries, minutes for complex analytical queries on large datasets
Write Performance Profile
Throughput: Tens of thousands of records/sec to hundreds of MB/s, depending on cluster size and write pattern (append vs. upsert)
Payload Size: Typically MBs to GBs per transaction/batch
Processing Latency: Seconds to minutes for batch writes, potentially sub-second for micro-batches or streaming appends
Architecture Diagram
#architecture_diagramReference diagram of the technology's internal architecture.
Component Architecture & Topology
Features
#featuresCatalogued product capabilities and what each one does.
Empirical data for features is currently being compiled in the global catalog.
Typical Use Cases
#typical_use_casesScenarios the technology is commonly chosen for.
Empirical data for typical use cases is currently being compiled in the global catalog.
Known Customers
#known_customersPublicly referenced organisations using the technology.
Empirical data for known customers is currently being compiled in the global catalog.
Known Integrations
#known_integrationsOther products and services it is documented to work with.
Empirical data for known integrations is currently being compiled in the global catalog.
Connectors
#connectorsDirectional data connections to other technologies, with direction and maturity.
Empirical data for connectors is currently being compiled in the global catalog.
Reference Architectures
#reference_architecturesPublished architectures where the technology is used or mentioned.
No published reference architecture blueprints currently link to Delta Lake.
Security Features
#security_featuresBuilt-in security and access-control capabilities.
Empirical data for security features is currently being compiled in the global catalog.
Known Issues
#known_issuesDocumented limitations, defects and operational pitfalls.
Empirical data for known issues is currently being compiled in the global catalog.
Guidelines
#guidelinesRecommended practices for adopting and operating the technology.
Empirical data for guidelines is currently being compiled in the global catalog.
Standards & Compliance
#standards_and_complianceStandards, certifications and control requirements it maps to.
Empirical data for standards & compliance is currently being compiled in the global catalog.
Sources
#sourcesDocumentation and research references behind the recorded information.
Empirical data for sources is currently being compiled in the global catalog.
Expert Validation
#expert_validationWhether domain experts reviewed and confirmed the content.
Empirical data for expert validation is currently being compiled in the global catalog.