ChistaDATA · Real-time analytics · Updated September 2026

From Batch Processing to Real-Time Analytics with ClickHouse: Architecture, Migration and Proof

Moving from batch processing to real-time analytics means replacing the nightly ETL window with a tier that ingests continuously, becomes queryable in seconds and serves hundreds of concurrent users. ClickHouse was built for exactly that workload profile.

This engineering guide from ChistaDATA covers what changes when batch processing gives way to real-time analytics, the reference architecture we deploy, the MergeTree data model that keeps queries sub-second, streaming ingestion from Kafka and CDC, a six-phase reversible migration off a batch warehouse, and the system-table telemetry that proves the result. Every claim on the page maps to a mechanism you can inspect, not a slogan.

01

What changes

  • Hours of freshness become seconds
  • Few large loads become continuous inserts
  • Tens of analysts become thousands of users
  • Load windows give way to background merges

02

How it is built

  • Kafka, Debezium CDC and async inserts
  • Raw ReplicatedMergeTree, immutable
  • Incremental materialized views to rollups
  • TTL tiering to S3, Keeper for coordination

03

What is current in 2026

  • Async inserts on by default since 26.3 LTS
  • Lightweight UPDATE with patch parts (25.7+)
  • Iceberg and Delta writes, open catalogs
  • ClickHouse 26.8 LTS distributed optimizer

04

How it is proven

  • End-to-end freshness in seconds
  • p95 and p99 from system.query_log
  • Parts and merge backlog from system.parts
  • Parity checks against the old warehouse
SecondsEvent-to-query freshness on a well-built real-time pipeline, against hours for batchDesign target
11M rows/sInsert rate Cloudflare publicly described on its ClickHouse analytics clusterCloudflare engineering blog
1,000,000,000Rows aggregated in 3.34 s on 2 vCPUs in our own ClickHouse labChistaDATA lab, measured
15 minSeverity 1 response, 24×7×365, once the platform is in productionChistaDATA SLA
100%Open-source ClickHouse, no proprietary forks or lock-inApache 2.0

The shift

Batch processing vs real-time analytics: what actually changes

Batch processing collects data over a window, hourly, nightly or weekly, then transforms and loads it in scheduled jobs. The design optimises throughput per job, not time to insight. Every downstream consumer, from fraud rules to executive dashboards, inherits the schedule’s latency.

Real-time analytics inverts the contract. Events are ingested continuously, become queryable within seconds, and are served to many concurrent users and applications against a data set that is always current. The workload profile differs in every dimension: many small inserts instead of a few large loads, continuous background merging instead of a maintenance window, and interactive concurrency instead of a handful of long reports.

The engine therefore has to be built for that profile. Bolting streaming onto a batch warehouse produces micro-batches with warehouse-scale cost and warehouse-scale latency. ClickHouse was designed from the start for continuous ingestion and sub-second aggregation over columnar data, which is why it has become the default open-source engine when teams move from batch processing to real-time analytics.

DimensionBatch warehouseReal-time analytics on ClickHouse
Data freshnessHours to a daySeconds
Ingestion patternFew large loadsContinuous streams and server-side batched inserts
Query concurrencyTens of analystsHundreds to thousands of users and APIs
Typical p95 latencySeconds to minutesSub-second on pre-aggregated paths
Cost driverCompute reserved for load windowsCompression ratio and merge efficiency
Failure surfaceMissed load windowConsumer lag, merge backlog, replication lag

The problem

Where batch processing architectures break down

None of these are performance bugs. They are consequences of a design that assumed data could wait.

Latency is a blind spot, not a delay

A fraud rule that runs on last night’s data cannot stop this morning’s transaction. An SRE dashboard refreshed hourly hides a ten-minute outage. The cost of batch processing is every decision made without data that already existed.

Concurrency was never in the design

Batch warehouses assume a small number of heavy queries. Exposing them to customer-facing dashboards or product analytics creates queue depth, admission-control throttling and unpredictable p99 latency.

Streaming sources arrive anyway

Kafka topics, CDC from PostgreSQL and MySQL, clickstreams and telemetry are already continuous. Forcing them into hourly loads adds staging layers, duplicate handling and late-arrival logic that the analytics engine should own natively.

Cost scales with the wrong variable

Reserved compute sized for the nightly peak sits idle the rest of the day. Real-time engines spread work continuously and win on compression, so cost tracks data volume rather than load-window size.

Batch processing versus real-time analytics timeline: nightly ETL delivers data hours later, while streaming ingestion into ClickHouse makes events queryable in seconds
Figure 1. The same event under batch processing and under real-time analytics. Batch adds staging, a scheduled ETL window and a load before anyone can query; the streaming path makes the row visible seconds after it is produced.

The engine

Why ClickHouse is the engine for real-time analytics

ClickHouse’s advantage is not one feature but four properties that a batch-to-real-time migration needs at the same time. Each maps to a mechanism you can inspect in system.* tables.

1. Query performance at scale

Columnar storage in the MergeTree family, a sparse primary index with configurable granularity, vectorized execution and codec-level compression (LZ4, ZSTD, Delta, DoubleDelta, Gorilla) let aggregations scan billions of rows with a fraction of the I/O. We explain the execution model in ClickHouse vectorized query processing and the storage side in columnar vs row-based databases, measured on 1 billion rows.

2. Native streaming ingestion

The Kafka table engine, the official ClickHouse Kafka Connect sink with exactly-once delivery, asynchronous inserts (on by default since 26.3 LTS) and Debezium-based CDC deliver rows to MergeTree parts continuously, without an external micro-batch scheduler.

3. High concurrency for interactive applications

Incremental materialized views, projections, the query cache, workload-scoped settings profiles and, since 26.8 LTS, a cost-based distributed optimizer keep p95 predictable when hundreds of dashboards hit the same tables.

4. Open ecosystem, open licence

Apache 2.0 licensed, deployable on bare metal, Kubernetes or any cloud, with S3 tiering for cold data, Apache Iceberg and Delta Lake reads and writes for the lakehouse, and first-class drivers for Grafana, Superset, Metabase and every major language.

Reference architecture

Ingest, transform, serve: the real-time analytics pipeline on ClickHouse

This is the architecture ChistaDATA deploys for batch-to-real-time programmes. It is stateless above ClickHouse: replay a Kafka offset range or re-snapshot a CDC source and the same materialized views rebuild the same rollups. That property is what makes the cutover reversible.

Real-time analytics reference architecture on ClickHouse: Kafka, Debezium CDC and application inserts feed a raw ReplicatedMergeTree table, incremental materialized views maintain rollups, and dashboards, APIs and alerting read through a proxy layer
Figure 2. The three planes of a real-time analytics pipeline on ClickHouse: ingestion, transformation at write time, and serving from rollups, with Keeper, S3 tiering and telemetry underneath.

01 · INGEST

Streams land as parts

Kafka, Redpanda or Kinesis topics, Debezium CDC from PostgreSQL and MySQL, and application inserts over HTTP or native protocol. Target tens of thousands of rows per insert, or let async inserts coalesce small writes server-side.

02 · TRANSFORM

Once, at write time

A raw MergeTree table holds immutable events. Incremental materialized views fan out on insert into SummingMergeTree or AggregatingMergeTree rollups by minute, hour, tenant or dimension set. No per-query transformation.

03 · SERVE

Rollups, not raw scans

Dashboards, embedded product analytics, alerting and ML feature reads query the rollups through chproxy or a load balancer, with per-workload quotas and settings profiles. Cold partitions tier to object storage under TTL.

04 · OPERATE

Keeper, tiering, telemetry

ClickHouse Keeper coordinates replication, TTL moves hot data to S3-backed disks after its useful window, and every plane exports system-table metrics to Prometheus and Grafana.

Data modelling

MergeTree data modelling for freshness and query latency

The single most important decision in a real-time analytics deployment is the ORDER BY key of the raw table, because it defines both the sparse primary index and the physical sort order inside each part. Lead with the low-cardinality columns every query filters on, then the time column. Partition by a coarse time bucket so that TTL moves and drops operate on whole partitions and inserts do not spray parts across many partitions.

The pattern below is minimal but production-shaped: a replicated raw events table, an hourly rollup and the materialized view that keeps the rollup current on every insert. Engine parameters are explicit by ChistaDATA convention. It is written for ClickHouse 25.8 LTS and later and verified on 26.8 LTS.

Design rule. Raw events are append-only and immutable; every business question is answered from a rollup or projection that a materialized view maintains. Mutations on the hot path are a design smell. Where late corrections are unavoidable, use the lightweight UPDATE available since 25.7, which writes patch parts instead of rewriting the partition.
-- Raw immutable events
CREATE TABLE IF NOT EXISTS analytics.events_raw ON CLUSTER '{cluster}'
(
    event_time   DateTime64(3, 'UTC') CODEC(Delta(8), ZSTD(1)),
    tenant_id    UInt32,
    event_type   LowCardinality(String),
    user_id      UInt64,
    amount       Decimal(18, 4),
    attrs        Map(String, String)
)
ENGINE = ReplicatedMergeTree('/clickhouse/tables/{shard}/analytics/events_raw', '{replica}')
PARTITION BY toYYYYMMDD(event_time)
ORDER BY (tenant_id, event_type, event_time)
TTL toDateTime(event_time) + INTERVAL 90 DAY TO VOLUME 'cold'
SETTINGS storage_policy = 'hot_cold',
         index_granularity = 8192;

-- Hourly rollup served to dashboards
CREATE TABLE IF NOT EXISTS analytics.events_1h ON CLUSTER '{cluster}'
(
    hour         DateTime('UTC'),
    tenant_id    UInt32,
    event_type   LowCardinality(String),
    events       UInt64,
    amount_sum   Decimal(18, 4),
    users        AggregateFunction(uniq, UInt64)
)
ENGINE = ReplicatedAggregatingMergeTree('/clickhouse/tables/{shard}/analytics/events_1h', '{replica}')
PARTITION BY toYYYYMM(hour)
ORDER BY (tenant_id, event_type, hour);

-- Maintained on every insert into events_raw
CREATE MATERIALIZED VIEW IF NOT EXISTS analytics.mv_events_1h ON CLUSTER '{cluster}'
TO analytics.events_1h
AS
SELECT
    toStartOfHour(event_time) AS hour,
    tenant_id,
    event_type,
    count()          AS events,
    sum(amount)      AS amount_sum,
    uniqState(user_id) AS users
FROM analytics.events_raw
GROUP BY hour, tenant_id, event_type;

Freshness is then a function of three things you can measure: Kafka consumer lag, the insert-to-part latency reported in system.part_log, and replication delay in system.replicas. Query latency on the serving path is a function of how much the rollup reduces the scanned row count, visible directly as read_rows in system.query_log.

Two additions belong in a 2026 design. Projections let a single raw table serve a second sort order or a pre-aggregation without a separate view. Refreshable materialized views, production-ready since 25.1, cover the small set of rollups that must be recomputed on a schedule rather than incrementally, for example a daily top-N that joins to a dimension table.

Current as of September 2026

What has changed for real-time analytics on ClickHouse recently

A batch-to-real-time design written against 23.x ClickHouse leaves capability on the table. These are the changes that alter the architecture or the operations of the pipeline above.

ChangeSinceWhat it means for a real-time pipeline
Async inserts enabled by default26.3 LTSSmall, frequent inserts are batched server-side without client changes. Review async_insert_max_data_size and async_insert_busy_timeout_ms: they now define the floor of your freshness budget.
Lightweight UPDATE with patch parts25.7+Late corrections no longer rewrite whole parts. CDC patterns that used ALTER TABLE ... UPDATE can move to patch parts; ReplacingMergeTree remains the choice for high-rate upserts.
Native JSON type25.3 GASemi-structured events can land as typed JSON with dynamic subcolumns instead of Map(String, String), with better compression and direct column access.
Iceberg and Delta Lake writes, catalog integrations25.8 to 26.8The cold tier can be an open lakehouse readable by Spark and Trino, not just S3-backed MergeTree parts. 26.8 LTS added Iceberg writes to S3 Tables and Puffin statistics.
Vector similarity index GA25.8 LTSReal-time feature stores and RAG retrieval can run beside the analytics tables instead of in a separate engine.
Cost-based distributed optimizer, parallel GROUP BY with adaptive merging26.8 LTSDistributed JOIN and aggregation plans change on upgrade. Re-baseline the serving queries with EXPLAIN PIPELINE before and after.
Parquet lazy materialisation26.8 LTSBackfills from object storage read far fewer bytes; ClickHouse reports around 77% fewer S3 reads on its own benchmarks, so history loads during migration are cheaper.
Release coverage: ChistaDATA delivers consulting, 24×7 support, managed services and remote DBA on ClickHouse 26.8 LTS (current long-term-support release, August 2026, supported to August 2027) and 26.3 LTS (supported to March 2027), plus every supported stable release between them, with upgrade engineering for 24.x and 25.x estates. Last verified September 2026.

Migration path

Migrating from batch processing to real-time analytics in six reversible phases

ChistaDATA runs batch-to-real-time migrations as staged, reversible programmes. The source system keeps serving until ClickHouse has proven parity on both data and latency, and every phase has an explicit rollback: stop the read cutover and the batch warehouse is still authoritative.

Six-phase migration from batch processing to real-time analytics on ClickHouse: assess, model, dual-write, validate, cut over reads, operate, each with a rollback gate
Figure 3. The six phases, each with an exit criterion and a rollback. The batch warehouse stays authoritative until the last consumer moves and a full reporting cycle passes clean.
01

Assess

Inventory queries, SLAs and consumers. Classify workloads by required freshness and concurrency. Identify which streams (Kafka, CDC, files) already exist and which batch jobs must become streams.

Exit criterionWorkload register with freshness class, concurrency and owner for every consumer.
02

Model

Design sort keys, partitions, rollups and TTL tiers from the actual query inventory. Size shards, replicas and ClickHouse Keeper. Define RPO and RTO and the replication topology before the first byte is loaded.

Exit criterionSchema, capacity model and DR design reviewed and signed off.
03

Dual-write

Attach ClickHouse as a second consumer of the same Kafka topics or a Debezium CDC stream. Backfill history from the warehouse in partition-sized batches, from Parquet on object storage where possible. Nothing downstream changes yet.

Exit criterionLive stream and backfill both landing; consumer lag inside target.
04

Validate

Run parity checks (row counts, sums, distinct counts by partition) and latency comparisons against production query mixes. Load-test concurrency with the real dashboard fleet, not synthetic users.

Exit criterionParity report and latency report from system.query_log, both inside SLO.
05

Cut over reads

Move consumers workload by workload, starting with the highest-freshness need. Keep the batch warehouse warm until the last consumer moves and one full reporting cycle passes clean. Rollback is a routing change.

Exit criterionAll consumers on ClickHouse; one clean reporting cycle; warehouse decommission approved.
06

Operate

SLOs and error budgets per workload, runbooks, quarterly restore and failover drills, upgrade management across LTS lines and capacity planning, under ChistaDATA managed services or your own team with our 24×7 support behind it.

Exit criterionSLOs green for a quarter; first restore drill passed with measured RTO.
Ingestion strategyBest fitTrade-off to plan for
Dual-write from the applicationGreenfield event producers you controlConsistency between writes; prefer a broker in between
Kafka fan-out to ClickHouseExisting event bus with topic historyExactly-once needs the Kafka Connect sink or idempotent keys; the Kafka table engine is at-least-once
CDC via DebeziumOLTP tables in PostgreSQL or MySQLUpdates and deletes need ReplacingMergeTree with FINAL or dedup-at-read, or lightweight UPDATE on 25.7+
Gradual read migrationLarge BI estates with many dashboardsSemantic-layer differences: SQL dialect, NULL handling, time zones

Measurement

Proving the outcome: the telemetry that defines real-time analytics

A migration is finished when the numbers say so. ChistaDATA reports three metrics for every real-time analytics engagement, each sourced from a ClickHouse system table rather than an estimate.

Freshness budget for real-time analytics on ClickHouse: producer to Kafka lag, async insert flush, part creation, replica fetch and query, with the system table that measures each stage
Figure 4. The freshness budget. Each stage between the event and the query adds latency, and each has a system table or metric that measures it. The sum is the number you report.

End-to-end freshness

The gap between the event timestamp and the moment the row is visible to SELECT. Compare now() to max(event_time) on the serving table, and correlate with Kafka consumer lag and insert events in system.part_log.

Serving latency

p95 and p99 of query_duration_ms from system.query_log per dashboard or API user, with read_rows and read_bytes to show how much the data model reduced scan volume. See the p99 latency playbook.

Ingestion health

Parts per partition and merge backlog from system.parts and system.merges; too-many-parts errors are the classic symptom of under-batched inserts. Replication delay from system.replicas bounds the freshness of read replicas.

Illustrative, not measured for you

A well-modelled hourly rollup typically reduces read_rows by two to four orders of magnitude versus scanning raw events, which is where sub-second p95 comes from. Actual ratios depend on cardinality and query shape; we measure them per workload during validation.

The three queries behind the report

-- Serving latency by dashboard user, last 24 h
SELECT
    user,
    count() AS queries,
    quantile(0.95)(query_duration_ms) AS p95_ms,
    quantile(0.99)(query_duration_ms) AS p99_ms,
    avg(read_rows) AS avg_read_rows
FROM system.query_log
WHERE type = 'QueryFinish'
  AND event_time >= now() - INTERVAL 24 HOUR
  AND query_kind = 'Select'
GROUP BY user
ORDER BY p99_ms DESC;

-- Ingestion health: parts per active partition
SELECT
    database,
    table,
    partition,
    count() AS active_parts,
    sum(rows) AS rows,
    max(modification_time) AS last_part_written
FROM system.parts
WHERE active
  AND database = 'analytics'
GROUP BY database, table, partition
ORDER BY active_parts DESC
LIMIT 20;

-- Freshness: how far behind is the serving table?
SELECT dateDiff('second', max(hour), now()) AS freshness_lag_seconds
FROM analytics.events_1h;

Use cases

Where moving from batch processing to real-time analytics pays back first

Fraud, risk and AML

Score transactions against rolling windows of behaviour while the transaction is still in flight. ClickHouse holds months of history at full granularity and answers window aggregations in milliseconds. Read: real-time AML on ClickHouse.

Observability and log analytics

Replace hourly log rollups and expensive Elasticsearch clusters with a single MergeTree tier: high-cardinality metrics, traces and logs queried together. Read: Elasticsearch to ClickHouse for observability.

Product and customer analytics

Embedded, customer-facing dashboards with per-tenant isolation, funnels, retention cohorts and A/B results updated as sessions happen, at concurrency levels a batch warehouse cannot serve.

Ad-tech, gaming and IoT telemetry

Billions of events a day from bidding, live game sessions or sensor fleets, with sub-second rollups for monetisation, matchmaking and anomaly alerting.

How ChistaDATA delivers

How ChistaDATA delivers batch-to-real-time on ClickHouse

ChistaDATA is a full-stack ClickHouse infrastructure operations company: consulting, 24×7×365 consultative support, platform engineering and managed services on 100% open-source ClickHouse. Every real-time analytics engagement is led by principal ClickHouse engineers and documented in decision-grade written deliverables.

  • Assessment and architecture. Query inventory, freshness classification, sort-key and partition design, shard, replica and Keeper sizing, RPO and RTO definition, and a migration roadmap with rollback at every phase. Start with a ClickHouse performance audit or ClickHouse consulting.
  • Proof of concept on your data. Representative workloads, real dashboard concurrency, parity and latency reports from system.query_log, and a cost model based on measured compression.
  • Production operations. 24×7 monitoring, SLOs and error budgets, runbook automation, upgrade management, quarterly restore and failover drills and capacity planning, through managed services, a remote ClickHouse DBA or our Data SRE practice.

Support is governed by an enterprise SLA: Severity 1 response in 15 minutes, Severity 2 in 12 hours, Severity 3 in 24 hours, Severity 4 in 48 hours. Standing caveat for every recommendation on this page: test before applying to production, and maintain a robust, drilled disaster-recovery posture.

FAQ

Batch processing to real-time analytics: frequently asked questions

What is the difference between batch processing and real-time analytics?

Batch processing transforms and loads data in scheduled windows, so every consumer sees data that is hours old. Real-time analytics ingests continuously and makes events queryable within seconds, serving many concurrent users against an always-current data set. The two differ in freshness, ingestion pattern, concurrency and cost model, not only in speed.

Why use ClickHouse for real-time analytics instead of a cloud data warehouse?

Cloud warehouses are built for large scheduled loads and a small number of heavy queries. ClickHouse is built for continuous inserts, background merges and high query concurrency, with native Kafka and CDC ingestion, incremental materialized views and columnar compression that keep both latency and cost predictable. It is also open source, so the same platform runs on any cloud or on premises.

How fresh can data be in ClickHouse?

Seconds, end to end, on a correctly built pipeline. The budget is the sum of producer-to-broker lag, the async insert flush interval, part creation and replica fetch. Each stage is measurable from Kafka consumer metrics, system.part_log and system.replicas, and each can be tuned down to sub-second on the hot path.

Can we migrate from batch processing to real-time analytics without downtime?

Yes. The six-phase path above dual-writes ClickHouse beside the existing warehouse, validates parity and latency, then moves consumers one workload at a time. The warehouse stays authoritative until the last consumer moves, so rollback at any phase is a routing change rather than a data recovery.

Which ClickHouse version should a new real-time platform run?

The current LTS line. Since 26.3 LTS async inserts are on by default, and 26.8 LTS brought the cost-based distributed optimizer and Iceberg writes. The release coverage line on this page is updated as each LTS ships, and we run upgrade engineering for estates on older lines.

How does ChistaDATA support real-time ClickHouse platforms in production?

Through 24×7×365 consultative support with a 15-minute Severity 1 response, fully managed services under contracted SLOs, or a named remote ClickHouse DBA inside your team. All three include monitoring from system tables, runbooks, quarterly restore drills and upgrade management.

Start the transformation

Move from batch processing to real-time analytics with ClickHouse

Bring your query inventory, your streams and your freshness targets. ChistaDATA engineers will map a staged, reversible migration and prove the result in your own telemetry.

All guidance on this page is general engineering advice. Test every change on a non-production environment before applying it to production, and maintain a robust backup and disaster recovery posture. Version claims are as published in September 2026.