LabHub

Blog

Streaming vs Batch Redefined: Flink, RisingWave, Materialize, CDC, Streaming SQL and the Pragmatism of Real Time (2025)

한국어English日本語中文

Season 5 Ep 2 — If Ep 1 was "where data is stored", Ep 2 is "how fast data flows". The real-time frenzy of 2020–2023, and the return to pragmatism in 2024–2025.

Prologue — "Real time is an option, not a default"

Every keynote at a data conference between 2019 and 2022 said "all data must become real time". The reality of 2025 is different:

The right answer in 2025:

"Divide data into freshness tiers based on the SLA of each dataset and metric, and use the tool that fits each tier."

This post makes those tiers and tool choices concrete.


Chapter 1 · Freshness Tiers

1.1 A Five-tier Framework

TierLatencyExampleTools
Real-timems–secondsTransaction monitoring, anomaly detection, fraudFlink, Kafka Streams
Near-real-time1–5 minOperational dashboards, alertsFlink, RisingWave, Materialize
Fresh5–60 minInventory, ad optimizationStreaming append + rollup
Daily24 hoursBI, reportsSpark, dbt, SQL warehouse
HistoricalWeekly/monthlyAnalysis, ML trainingBatch yearly/monthly

1.2 Mapping Each Metric and Table to a Tier

1.3 Decision Principles


Chapter 2 · The 2025 Version of Lambda and Kappa

2.1 Lambda (2014)

2.2 Kappa (2014, Jay Kreps)

2.3 2025: Unified on Lakehouse

2.4 The "Streaming + Materialized" Pattern


Chapter 3 · Comparing the Four Major Streaming Engines

3.2 Spark Structured Streaming

3.3 Kafka Streams / ksqlDB

3.4 RisingWave

3.5 Materialize

3.6 Comparison Table

EngineLatencyComplexitySQLUse in KoreaCharacteristics
Flinkms–secondsHighYesCommonIndustry standard, strong state management
Spark SSSeconds–minutesMediumYesVery commonDatabricks friendly
Kafka StreamsSecondsMediumksqlDBModerateBuilt into Kafka
RisingWaveSecondsLowPostgresGrowingSimple to operate, SaaS/OSS
MaterializeSecondsMediumYesRareStrong incremental views

3.7 Selection Guide


Chapter 4 · CDC (Change Data Capture)

4.1 Why CDC Is Central

4.2 Implementation Methods

4.3 Tools

4.4 The CDC → Iceberg Pattern

PostgresDebeziumKafkaFlinkIceberg

4.5 Practical Traps


Chapter 5 · Iceberg v3 and Real-time Upsert

5.1 Iceberg Version History

5.2 Two Kinds of Row-level Delete

5.3 The Real-time Upsert Workflow

  1. Flink reads the CDC events
  2. Generates an equality delete plus an insert based on the PK
  3. Iceberg reflects it in a snapshot
  4. Periodic compaction cleans up the delete files

5.4 Performance Cautions


Chapter 6 · The Rise of Streaming SQL

6.1 Why SQL

6.3 ksqlDB

6.4 The Postgres Compatibility of RisingWave

6.5 Materialize


Chapter 7 · The Cost and Latency Trade-off

7.1 Cost Components

7.2 Representative Cost Comparison (monthly, mid-size scale)

OptionMonthly costLatency
Batch (Airflow + Spark, daily)Low ($1–5k)24 hours
Micro-batch (5 min)Medium ($3–10k)5 min
Structured StreamingMedium–high ($5–20k)Seconds–minutes
Flink clusterHigh ($10–30k+)ms–seconds
Managed (RisingWave/Confluent)Medium–high ($7–25k)Seconds

7.3 Cost Reduction Techniques


Chapter 8 · Observability and Debugging

8.1 Key Metrics

8.2 Observability Tools

8.3 Debugging

8.4 Alerting


Chapter 9 · Failure, Recovery and SLA

9.1 Designing the SLA

9.2 Recovery Strategy

9.3 Reprocessing

9.4 Multi-region


Chapter 10 · Practical Patterns for Streaming plus Lakehouse

10.1 Streaming on Medallion

10.2 CDC → Silver

10.3 Event Sourcing

10.4 Real-time Feature Store

10.5 Streaming ETL Pipeline


Chapter 11 · Three Real-world Cases

11.1 An E-commerce Order Pipeline

11.2 Financial Transaction Monitoring

11.3 Game Telemetry


Chapter 12 · Streaming at Korean Companies

12.1 Traditional Patterns

12.3 Regulatory Considerations

12.4 Obstacles


Chapter 13 · Ten Anti-patterns

13.1 "Everything in Real Time"

Streaming even the tables you do not need, so cost and complexity explode.

13.2 Blind Faith in Exactly-once

Guaranteeing end-to-end exactly-once on both the source and the sink is not easy. Idempotent design is mandatory.

13.3 Skipping the Initial CDC Snapshot

Gaps appear and accuracy drops.

13.4 No Automatic Propagation of Schema Changes

Downstream pipelines break.

13.5 Kafka Retention Too Short

Reprocessing becomes impossible.

13.6 Checkpoint Interval Too Long

Recovery cost and the amount to reprocess explode when something fails.

13.7 Keeping State Without Limits

Flink keyed state grows without bound, leading to OOM.

13.8 No Compaction of Delete Files

Iceberg read performance degrades.

13.9 Under-allocating Memory and CPU

A chain reaction of back-pressure.

13.10 No Observability or Alerting

Customers discover the incident before you do.


Chapter 14 · Checklist — Twelve Points Before Launching Streaming


Chapter 15 · Next Up — Season 5 Ep 3: "Comparing OLAP Engines in 2025"

Now that streaming and batch share the same storage, the next question is "who queries fastest on top of it?".

Things get genuinely interesting only after you accept the 2025 reality that "a single engine cannot do everything".

See you in the next post.


Summary: Streaming in 2025 was redefined from "everything in real time" to "freshness tiers based on SLA". You place Flink, Spark SS, RisingWave, Materialize and ksqlDB across the five tiers of Real-time, Near-real-time, Fresh, Daily and Historical, use CDC to flow changes from the operational DB into the Lakehouse, and handle real-time upserts with Iceberg v3 row-level delete. "Unified on Lakehouse" rather than Lambda/Kappa is the dominant pattern, and the trade-offs of cost, latency and complexity are designed deliberately. "Real time is an option, not a default" — that is the pragmatism of 2025.

Comments

No comments yet.

Sign in to leave a comment