Engineering posts about Data Lake

Curated summaries and key learnings for engineers working with Data Lake.

Databricks
8m

How Indra unified EV charging data on Databricks

Indra Renewable Technologies faced significant challenges with fragmented data systems as their electric vehicle (EV) charging operations expanded. The company consolidated its data management by...

Databricks
43m

Vertical Advantage: Transforming Industries with Lakebase and Agentic AI

The article highlights how Databricks Lakebase transforms industry operations by providing a unified platform for data analytics and AI. It emphasizes the integration of operational data with...

Databricks
4m

Building for the AI Era: Lakebase, Streaming, and Lakehouse Innovations at VLDB 2026

The article highlights key innovations presented at VLDB 2026 by Databricks, focusing on the evolving landscape of database engineering in the AI era. It introduces Lakebase, a serverless PostgreSQL...

Databricks
5m

Introducing Governance Hub: Intelligent, account-level governance over your Databricks estate

The Governance Hub is a centralized platform designed to enhance governance over Databricks estates by providing insights into data health, AI usage, and cost management across multiple cloud...

Databricks
16m

Open Table Formats Explained: Iceberg vs. Delta vs. Hudi

The article explores the concept of open table formats that enhance data lakes by providing features such as ACID transactions, schema evolution, and time travel. It compares three primary formats:...

Databricks
19m

Data Mesh vs. Data Fabric: Key Differences and How the Lakehouse Resolves the Debate

The article contrasts data mesh and data fabric, emphasizing their distinct approaches to data ownership and governance. Data mesh advocates for decentralized ownership by domain teams, treating data...

Databricks
5m

Run, debug, and scale Databricks workloads from your local IDE

The article discusses the latest enhancements in Databricks that allow developers to run, debug, and scale workloads directly from local IDEs like VS Code and Cursor. It emphasizes the importance of...

Netflix
11m

A Tale of Two Flink Autoscalers

The article explores Netflix's experience with two autoscalers for Apache Flink, highlighting the evolution from an in-house solution to adopting an open-source autoscaler from the Flink community....

AWS
4m

AWS Glue 6.0 now available with 30% lower price and full Apache Iceberg v3 support

AWS Glue 6.0 has been released with significant enhancements, including a 30% price reduction and full support for Apache Iceberg v3. This version is built on Apache Spark 4.1, Python 3.12, and Scala...

Dropbox
11m

Improving infrastructure efficiency for growing demand in the age of AI

The article outlines Dropbox's approach to enhancing infrastructure efficiency amidst growing AI demands. It emphasizes the importance of a system-level perspective in optimizing various...

Databricks
11m

How Databricks Feature Store serves features with sub-second freshness

The article discusses the capabilities of Databricks Feature Store in delivering features with sub-second freshness, essential for real-time machine learning applications such as fraud detection and...

Databricks
7m

Using AI_Functions in Your Data Warehouse: Top Use Cases

The article explores the utilization of AI Functions in data warehouses, emphasizing their role in processing both structured and unstructured data. It highlights the inefficiencies of traditional...

Databricks
9m

How a major freight railroad scaled pipeline creation with Genie Code

The article details how a major freight railroad utilized Databricks Genie Code to automate the modernization of its data pipelines, significantly reducing the time required for pipeline delivery...

Databricks
4m

How Amtrak is building the data backbone for its largest transformation in over 50 years

Amtrak is undergoing a significant transformation by building a digital intelligence platform named Rail Intelligence, which integrates various data sources into a unified system. This platform...

Databricks
13m

Modern Risk Demands a Real-Time Foundation: The CRO’s Mandate

The article highlights the critical need for real-time data architecture in enterprise risk management, particularly for financial institutions facing volatile markets. It discusses how legacy...

Databricks
7m

Taking AUTO CDC to the next level: Solving the hardest real-world use cases

The article delves into the advancements of AUTO CDC (Change Data Capture) in Apache Spark, particularly addressing complex real-world use cases that traditional CDC patterns struggle to handle. It...

Databricks
9m

Introducing FILE type: a native column type for multimodal data

The article introduces the FILE type, a new column type designed to store unstructured data as a native column in databases, allowing for unified governance and management alongside structured data....

Salesforce
5m

How Standardizing Product Telemetry Reduced Time to Insight by 97%

The article outlines how Salesforce standardized its product telemetry through the implementation of a Product Data Platform (PDP), which significantly reduced the time to insight for product...

Databricks
4m

Ingest semi-structured data faster and more efficiently with Variant - Now Generally Available

The article introduces Variant, a new data type designed to facilitate the ingestion of semi-structured data such as JSON, XML, and CSV into Databricks. It addresses the historical trade-off between...

Databricks
6m

Backstage with Lakebase, part 3

This article is the third installment in a series exploring the integration of Databricks Lakebase with Backstage to enhance operational visibility and cost management in cloud environments. It...