100% open source · 24×7×365 operations · zero vendor lock-in

ChistaDATA Server for ClickHouse: production ClickHouse, engineered and operated around the clock

Architecture · Performance engineering · HA/DR · Ingestion · Security · 24×7 support

ChistaDATA Server is how we run open-source ClickHouse for enterprises that cannot afford a slow dashboard or a lost partition. It is a 100% open-source, fully compatible ClickHouse distribution combined with the engineering blueprint, automation and 24×7×365 operations team that keep it fast, replicated and recoverable in production.

Your data stays in standard MergeTree formats, on your infrastructure, in your cloud account. If you ever leave, you leave with a working cluster.

At a glance

What it is
ChistaDATA Server for ClickHouse: open-source ClickHouse plus the ChistaDATA operating standard
Where it runs
On-premises, AWS, Azure, GCP, Kubernetes or hybrid
Severity 1 response
15 minutes, 24×7×365
ClickHouse versions
Release coverage: ChistaDATA delivers consulting, 24×7 support, managed services and remote DBA on ClickHouse 26.9 (latest stable release, September 2026), 26.8 LTS (current long-term-support release, August 2026, supported to August 2027) and 26.3 LTS (supported to March 2027), plus every supported stable release in between, with upgrade engineering for 24.x and 25.x estates, including 25.8 LTS, which reached end of support in August 2026. Last verified October 2026.
Lock-in
None. Standard ClickHouse formats, SQL and protocols

24×7×365Follow-the-sun ClickHouse operations from the San Francisco Bay Area and eleven global offices
15 minSeverity 1 response commitment. S2 in 12 hours, S3 in 24 hours, S4 in 48 hours
12 / yrUpstream ClickHouse releases a year, two of them LTS, each one tracked and risk-assessed before it reaches you
0Proprietary storage formats, licence keys or closed forks between your data and your team

Definition

What ChistaDATA Server for ClickHouse is, and what it is not

ClickHouse is one of the fastest open-source columnar databases for analytical workloads, with more than 49,000 GitHub stars and a release train of roughly one stable version a month and two long-term support (LTS) versions a year. Speed on a laptop is easy. Speed on a replicated, sharded, 24×7 production estate is an engineering discipline.

ChistaDATA Server packages that discipline. We announced it in January 2022 as a ClickHouse distribution that stays 100% open source and fully compatible with upstream ClickHouse. Today it is delivered as three layers:

  • The server. Upstream-compatible ClickHouse builds pinned to a supported LTS or stable line, with the settings, storage policies and Keeper topology we have validated for your workload.
  • The blueprint. Schema, sort-key, partitioning, replication and DR standards written as versioned runbooks, not tribal knowledge.
  • The operators. ClickHouse engineers on call 24×7×365, working from system tables and evidence, with a 15-minute Severity 1 response.

What it is not: a closed fork, a proprietary cloud, or a licence you cannot walk away from. Every table we build reads and writes on stock ClickHouse.

ChistaDATA Server for ClickHouse reference architecture with shards, ClickHouse Keeper, tiered storage and 24x7 operations
Figure 1. ChistaDATA Server reference architecture: routing, replicated shards, a Raft-based Keeper quorum, hot and cold storage tiers, and the operations layer on top.

Public evidence

The scale open-source ClickHouse is proven at

We do not publish invented benchmarks. The strongest argument for running ClickHouse seriously is what operators have already published about it. Cloudflare’s engineering team documented its HTTP analytics pipeline on ClickHouse in 2018, and the numbers still set the bar for what a well-engineered cluster can absorb.

6MHTTP requests per second on average, with peaks of 8M
11Mrows per second inserted across all pipelines
47 Gbpsaverage insertion bandwidth
36ClickHouse nodes with 3× replication

The storage outcome is the part most teams underestimate. Each request arrived as a 1,630-byte message, compressed to 360 bytes with zstd, and settled into ClickHouse at 36.74 bytes per row. That is roughly 44 times smaller than raw, across more than 100 columns per request and a 365-day retention requirement.

Those results come from sort-key design, codec choice, partitioning and merge behaviour, which is exactly the layer ChistaDATA Server standardises. Your own numbers will differ with your data, and we measure them on your workload before we promise anything.

Source: Cloudflare engineering blog, HTTP Analytics for 6M requests per second using ClickHouse. ClickHouse project statistics: ClickHouse on GitHub.

Cloudflare ClickHouse compression figures, 1,630 bytes raw to 36.74 bytes per row
Figure 2. Bytes per HTTP request at Cloudflare, as published by Cloudflare in March 2018. Public figures, not ChistaDATA measurements.

The operating standard

Eight engineering pillars of ChistaDATA Server

Each pillar has an owner, a runbook and a system table that proves it is working. If a pillar cannot be measured, it is not part of the standard.

01

Provisioning and lifecycle

Repeatable cluster builds on VMs, bare metal or Kubernetes, with LTS-pinned versions and rolling upgrades.

Evidence: system.clusters, version()

02

Ingestion engineering

Kafka, Debezium, Kinesis, S3 and Flink paths with batching and back-pressure designed in.

Evidence: system.asynchronous_inserts

03

Query performance

Sort keys, skip indexes, projections and materialized views chosen from real query logs.

Evidence: system.query_log

04

Scale-out

Sharding keys, resharding plans, parallel replicas and replica routing sized to growth.

Evidence: system.parts, system.disks

05

High availability and DR

ReplicatedMergeTree, Keeper quorum design, cross-region replicas and quarterly restore drills.

Evidence: system.replicas

06

Observability

Prometheus and Grafana dashboards, alerting on merges, mutations, replication and Keeper health.

Evidence: system.metrics, system.events

07

Security and governance

RBAC, row policies, quotas, settings profiles, TLS everywhere and auditable access.

Evidence: system.grants, system.session_log

08

Cost and FinOps

Right-sizing, codec tuning, TTL moves to object storage and query cost attribution.

Evidence: system.parts compression ratios

Measurement first

Performance engineering on ChistaDATA Server, measured in system tables

Every recommendation we make starts from telemetry the server already records. The first query we run on a new estate is a p95 and p99 latency profile by normalized query shape, because the slowest 1% of queries is where dashboards time out and where hardware budgets go.

-- p95 / p99 latency and read volume by query shape, last 24 hours
SELECT
    normalized_query_hash,
    any(query)                                   AS sample_query,
    count()                                      AS executions,
    quantile(0.95)(query_duration_ms)            AS p95_ms,
    quantile(0.99)(query_duration_ms)            AS p99_ms,
    formatReadableSize(sum(read_bytes))          AS total_read,
    formatReadableSize(max(memory_usage))        AS peak_memory
FROM system.query_log
WHERE type = 'QueryFinish'
  AND event_date >= today() - 1
  AND event_time >= now() - INTERVAL 24 HOUR
GROUP BY normalized_query_hash
ORDER BY p99_ms DESC
LIMIT 20;

The second check is part health. Too many active parts per partition is the most common root cause of TOO_MANY_PARTS errors, insert throttling and merge backlogs. ClickHouse starts delaying inserts at the parts_to_delay_insert threshold and rejects them at parts_to_throw_insert, so we alert well before either.

-- Active parts per partition, compared with the table's own thresholds
SELECT
    database,
    table,
    partition,
    count()                                      AS active_parts,
    formatReadableSize(sum(bytes_on_disk))       AS on_disk,
    round(sum(data_uncompressed_bytes) / sum(data_compressed_bytes), 1) AS compression_ratio
FROM system.parts
WHERE active
GROUP BY database, table, partition
ORDER BY active_parts DESC
LIMIT 20;
SignalSourceWhat ChistaDATA Server does with it
Query latency p95 / p99system.query_logRanks query shapes by tail latency, then tests sort-key, projection or materialized-view fixes against a replay of the same shapes.
Parts per partitionsystem.partsAlerts at half of parts_to_delay_insert, then fixes batching or partition granularity rather than raising limits.
Merge and mutation backlogsystem.merges, system.mutationsFlags stuck mutations and merge pressure before they turn into disk exhaustion or replication lag.
Replication queue and delaysystem.replicas, system.replication_queueTracks absolute delay per replica and queue size so failover targets are always current.
CPU hot spotssystem.trace_log, EXPLAIN PIPELINEProfiles where the pipeline spends time and whether parallelism is actually used.
Recommendations on this page are general engineering guidance. Test every change on a representative staging workload before applying it to production, and keep a verified backup and DR posture in place.

Streaming and CDC

Ingestion pipelines that keep ClickHouse fresh without drowning it

Most ClickHouse incidents we are called into start at ingestion, not at query time. Small, frequent inserts create too many parts, merges fall behind, and the cluster spends its CPU compacting instead of answering questions.

ChistaDATA Server standardises the path from source to table:

  • Apache Kafka through the Kafka table engine or sink connectors, with consumer lag tracked per partition.
  • Debezium CDC from PostgreSQL, MySQL and SQL Server into ReplacingMergeTree or CollapsingMergeTree designs that handle updates and deletes correctly.
  • Kinesis, RabbitMQ and Apache Flink through buffered sinks sized to the target batch.
  • S3 and GCS files through s3Queue and table functions for continuous file ingestion.
  • Asynchronous inserts where clients cannot batch, so the server does the buffering.

Then materialized views pre-aggregate the hot paths, so dashboards read thousands of rows instead of billions.

ClickHouse ingestion pipeline on ChistaDATA Server from Kafka, Debezium, Kinesis, S3 and Flink into MergeTree
Figure 3. Every source lands through a buffered, back-pressured path, then fans out to materialized views and serving tables.

Reliability engineering

High availability and disaster recovery with ChistaDATA Server

A replica you have never failed over to is a hope, not a plan. ChistaDATA Server treats availability as something you measure: replication delay, Keeper quorum health, backup age and the time a restore actually takes.

ClickHouse Keeper uses Raft, so quorum arithmetic is fixed: a 3-voter ensemble tolerates one failure and a 5-voter ensemble tolerates two. We place voters across failure domains, keep Keeper off the busiest data nodes, and never run an even number of voters.

For regional DR we combine asynchronous replicas in a second region with checksummed full and incremental backups in object storage, and we run restore and failover drills every quarter with the customer’s team watching.

ChistaDATA Server ClickHouse high availability and disaster recovery two-region topology with Keeper quorum
Figure 4. Two-region pattern: replicated shards and a Keeper quorum in each region, backups in object storage, drills every quarter.

Every table is created with an explicit engine definition. We never rely on defaults in customer estates, because defaults change between releases.

CREATE TABLE analytics.events ON CLUSTER 'prod_cluster'
(
    event_date   Date,
    event_time   DateTime64(3, 'UTC'),
    tenant_id    UInt32,
    event_type   LowCardinality(String),
    user_id      UInt64,
    payload      String CODEC(ZSTD(3))
)
ENGINE = ReplicatedMergeTree('/clickhouse/tables/{shard}/analytics/events', '{replica}')
PARTITION BY toYYYYMM(event_date)
ORDER BY (tenant_id, event_type, event_time)
TTL event_date + INTERVAL 90 DAY TO VOLUME 'cold'
SETTINGS index_granularity = 8192, storage_policy = 'hot_cold';
Failure scenarioChistaDATA Server designHow recovery is proven
Single node lossTwo or more replicas per shard, load balancing away from unhealthy replicasReplica kill test; replication delay back to baseline
Keeper voter loss3 or 5 voters across failure domainsVoter stop test; inserts continue, system.zookeeper_connection healthy
Zone or data-centre lossReplicas spread across zones; DR replicas in a second regionQuarterly failover drill with measured RTO
Logical corruption or bad DDLPoint-in-time backups, retention matched to the audit windowQuarterly restore to an isolated cluster with row-count validation

Security and compliance posture

Security and governance built into every cluster

ChistaDATA Server ships with access control configured, not bolted on. We use SQL-driven RBAC, row policies for multi-tenant isolation, quotas and settings profiles to stop runaway queries, and TLS on client and inter-server traffic.

Authentication integrates with LDAP and Kerberos where your estate requires it, and every login and query is auditable through system.session_log and system.query_log.

  • Least-privilege roles for applications, analysts and operators
  • Row and column restrictions for regulated or multi-tenant data
  • Encryption in transit, and encrypted disks for data at rest
  • Credentials held as secrets, never in configuration files
  • Controls mapped to GDPR, HIPAA, SOX, PCI DSS and SOC 2 requirements, with your auditors owning the certification

Vendor-neutral view

ChistaDATA Server vs DIY ClickHouse vs a managed cloud service

We are vendor-neutral by principle, and we will tell you when a different model fits better. A fully managed ClickHouse service can be the right answer for small teams with no residency constraints. ChistaDATA Server is built for estates that need control of the infrastructure and a team accountable for the outcome.

DimensionDIY open-source ClickHouseFully managed cloud serviceChistaDATA Server
Where data livesYour infrastructureProvider’s infrastructure or tenancyYour infrastructure or your cloud account
EngineOpen-source ClickHouseProvider build; some features differ from open source100% open-source ClickHouse, upstream compatible
UpgradesYour teamProvider scheduleLTS-pinned, risk-assessed, rolling, on your schedule
24×7 expertiseOnly if you hire itPlatform support; schema and query design usually yoursClickHouse engineers on call, S1 in 15 minutes
Exit costNoneMigration projectNone; the cluster stays yours

Reference patterns

Workload patterns ChistaDATA Server is engineered for

These are reference designs we use as starting points, not customer stories. Each one begins with a measured assessment of the actual data and queries.

Financial services

Real-time fraud and risk analytics

Sub-second aggregation over transaction streams from Kafka, with ReplacingMergeTree for late updates and strict row-level access.

Digital advertising

Impression and attribution analytics

High-cardinality event tables, materialized rollups per campaign and multi-region replicas for client-facing dashboards.

Observability

Logs, metrics and traces

Wide event tables with aggressive codecs, TTL to object storage and query quotas that protect the cluster during incidents.

IoT and telemetry

Industrial time-series

Per-device sort keys, downsampling views and retention tiers that keep years of history at a fraction of hot-storage cost.

Engagement model

How a ChistaDATA Server engagement runs

01 · ASSESS

Measure the estate

Query-log profile, part health, Keeper and replication review, capacity and cost baseline. Delivered as a written report. See our ClickHouse performance audit.

02 · ENGINEER

Fix the foundations

Schema and sort-key redesign, ingestion batching, HA/DR topology and security baseline, staged and reversible. See ClickHouse consulting.

03 · OPERATE

Run it 24×7×365

Monitoring, on-call, upgrades, backups and drills under an SLA. See ClickHouse managed services and ClickHouse DBA.

04 · OPTIMISE

Keep improving

Quarterly reviews of p99 latency, cost per query and growth, with a roadmap for the next two quarters. See Data SRE.

Moving from Redshift, Snowflake, BigQuery, Druid, Pinot, Vertica or Elasticsearch? Our ClickHouse migration practice plans schema translation, dual-running and cutover with a rollback path at every step.

Commitments

ClickHouse support tiers and response SLAs

ChistaDATA Server customers are covered by the same enterprise severity model as our ClickHouse support contracts. Response clocks run 24×7×365, including weekends and public holidays.

S1Critical

15 minInitial response

Production down or data at risk

Typical exampleKeeper quorum lost, inserts rejected cluster-wide

S2High

12 hoursInitial response

Severe degradation, workaround possible

Typical exampleOne shard lagging, p99 latency several times baseline

S3Medium

24 hoursInitial response

Non-critical defect or performance question

Typical exampleSlow report query, merge tuning request

S4Low

48 hoursInitial response

Advice, planning or documentation

Typical exampleUpgrade planning, schema review

FAQ

ChistaDATA Server for ClickHouse: frequently asked questions

Is ChistaDATA Server a fork of ClickHouse?

ChistaDATA Server is 100% open source and fully compatible with upstream ClickHouse. Tables, SQL, client drivers and wire protocols behave as they do on stock ClickHouse, so you can move between ChistaDATA Server and community builds without data conversion.

Which ClickHouse versions do you support?

We support the current LTS and stable lines and plan upgrades against the upstream calendar of two LTS releases a year, each supported for a year. Release coverage: ChistaDATA delivers consulting, 24×7 support, managed services and remote DBA on ClickHouse 26.9 (latest stable release, September 2026), 26.8 LTS (current long-term-support release, August 2026, supported to August 2027) and 26.3 LTS (supported to March 2027), plus every supported stable release in between, with upgrade engineering for 24.x and 25.x estates, including 25.8 LTS, which reached end of support in August 2026. Last verified October 2026.

Can ChistaDATA Server run on Kubernetes?

Yes. We run ClickHouse on Kubernetes with an operator, on virtual machines and on bare metal, across AWS, Azure, GCP and on-premises data centres. The choice follows your platform standards, not ours.

Do you replace ClickHouse Keeper with ZooKeeper?

New builds use ClickHouse Keeper by default. We still support existing ZooKeeper estates and plan migrations from ZooKeeper to Keeper as a staged, reversible change.

What does it cost?

Pricing depends on cluster size, coverage hours and the scope of engineering. Most engagements start with a fixed-scope assessment. Contact us for a proposal.

Will we be locked in?

No. Your data stays in standard ClickHouse formats on infrastructure you control, and every runbook we write for your estate is yours to keep.

Next step

Put ChistaDATA Server under your ClickHouse estate

Tell us about your cluster, your latency targets and your worst week of the last quarter. A ClickHouse engineer will reply with a measured view of what to fix first.