100% open source · 24×7×365 operations · zero vendor lock-in
ChistaDATA Server for ClickHouse: production ClickHouse, engineered and operated around the clock
Architecture · Performance engineering · HA/DR · Ingestion · Security · 24×7 support
ChistaDATA Server is how we run open-source ClickHouse for enterprises that cannot afford a slow dashboard or a lost partition. It is a 100% open-source, fully compatible ClickHouse distribution combined with the engineering blueprint, automation and 24×7×365 operations team that keep it fast, replicated and recoverable in production.
Your data stays in standard MergeTree formats, on your infrastructure, in your cloud account. If you ever leave, you leave with a working cluster.
At a glance
- What it is
- ChistaDATA Server for ClickHouse: open-source ClickHouse plus the ChistaDATA operating standard
- Where it runs
- On-premises, AWS, Azure, GCP, Kubernetes or hybrid
- Severity 1 response
- 15 minutes, 24×7×365
- ClickHouse versions
- Release coverage: ChistaDATA delivers consulting, 24×7 support, managed services and remote DBA on ClickHouse 26.9 (latest stable release, September 2026), 26.8 LTS (current long-term-support release, August 2026, supported to August 2027) and 26.3 LTS (supported to March 2027), plus every supported stable release in between, with upgrade engineering for 24.x and 25.x estates, including 25.8 LTS, which reached end of support in August 2026. Last verified October 2026.
- Lock-in
- None. Standard ClickHouse formats, SQL and protocols
Definition
What ChistaDATA Server for ClickHouse is, and what it is not
ClickHouse is one of the fastest open-source columnar databases for analytical workloads, with more than 49,000 GitHub stars and a release train of roughly one stable version a month and two long-term support (LTS) versions a year. Speed on a laptop is easy. Speed on a replicated, sharded, 24×7 production estate is an engineering discipline.
ChistaDATA Server packages that discipline. We announced it in January 2022 as a ClickHouse distribution that stays 100% open source and fully compatible with upstream ClickHouse. Today it is delivered as three layers:
- The server. Upstream-compatible ClickHouse builds pinned to a supported LTS or stable line, with the settings, storage policies and Keeper topology we have validated for your workload.
- The blueprint. Schema, sort-key, partitioning, replication and DR standards written as versioned runbooks, not tribal knowledge.
- The operators. ClickHouse engineers on call 24×7×365, working from system tables and evidence, with a 15-minute Severity 1 response.
What it is not: a closed fork, a proprietary cloud, or a licence you cannot walk away from. Every table we build reads and writes on stock ClickHouse.
Public evidence
The scale open-source ClickHouse is proven at
We do not publish invented benchmarks. The strongest argument for running ClickHouse seriously is what operators have already published about it. Cloudflare’s engineering team documented its HTTP analytics pipeline on ClickHouse in 2018, and the numbers still set the bar for what a well-engineered cluster can absorb.
The storage outcome is the part most teams underestimate. Each request arrived as a 1,630-byte message, compressed to 360 bytes with zstd, and settled into ClickHouse at 36.74 bytes per row. That is roughly 44 times smaller than raw, across more than 100 columns per request and a 365-day retention requirement.
Those results come from sort-key design, codec choice, partitioning and merge behaviour, which is exactly the layer ChistaDATA Server standardises. Your own numbers will differ with your data, and we measure them on your workload before we promise anything.
Source: Cloudflare engineering blog, HTTP Analytics for 6M requests per second using ClickHouse. ClickHouse project statistics: ClickHouse on GitHub.
The operating standard
Eight engineering pillars of ChistaDATA Server
Each pillar has an owner, a runbook and a system table that proves it is working. If a pillar cannot be measured, it is not part of the standard.
Provisioning and lifecycle
Repeatable cluster builds on VMs, bare metal or Kubernetes, with LTS-pinned versions and rolling upgrades.
Evidence: system.clusters, version()
Ingestion engineering
Kafka, Debezium, Kinesis, S3 and Flink paths with batching and back-pressure designed in.
Evidence: system.asynchronous_inserts
Query performance
Sort keys, skip indexes, projections and materialized views chosen from real query logs.
Evidence: system.query_log
Scale-out
Sharding keys, resharding plans, parallel replicas and replica routing sized to growth.
Evidence: system.parts, system.disks
High availability and DR
ReplicatedMergeTree, Keeper quorum design, cross-region replicas and quarterly restore drills.
Evidence: system.replicas
Observability
Prometheus and Grafana dashboards, alerting on merges, mutations, replication and Keeper health.
Evidence: system.metrics, system.events
Security and governance
RBAC, row policies, quotas, settings profiles, TLS everywhere and auditable access.
Evidence: system.grants, system.session_log
Cost and FinOps
Right-sizing, codec tuning, TTL moves to object storage and query cost attribution.
Evidence: system.parts compression ratios
Measurement first
Performance engineering on ChistaDATA Server, measured in system tables
Every recommendation we make starts from telemetry the server already records. The first query we run on a new estate is a p95 and p99 latency profile by normalized query shape, because the slowest 1% of queries is where dashboards time out and where hardware budgets go.
-- p95 / p99 latency and read volume by query shape, last 24 hours
SELECT
normalized_query_hash,
any(query) AS sample_query,
count() AS executions,
quantile(0.95)(query_duration_ms) AS p95_ms,
quantile(0.99)(query_duration_ms) AS p99_ms,
formatReadableSize(sum(read_bytes)) AS total_read,
formatReadableSize(max(memory_usage)) AS peak_memory
FROM system.query_log
WHERE type = 'QueryFinish'
AND event_date >= today() - 1
AND event_time >= now() - INTERVAL 24 HOUR
GROUP BY normalized_query_hash
ORDER BY p99_ms DESC
LIMIT 20;The second check is part health. Too many active parts per partition is the most common root cause of TOO_MANY_PARTS errors, insert throttling and merge backlogs. ClickHouse starts delaying inserts at the parts_to_delay_insert threshold and rejects them at parts_to_throw_insert, so we alert well before either.
-- Active parts per partition, compared with the table's own thresholds
SELECT
database,
table,
partition,
count() AS active_parts,
formatReadableSize(sum(bytes_on_disk)) AS on_disk,
round(sum(data_uncompressed_bytes) / sum(data_compressed_bytes), 1) AS compression_ratio
FROM system.parts
WHERE active
GROUP BY database, table, partition
ORDER BY active_parts DESC
LIMIT 20;| Signal | Source | What ChistaDATA Server does with it |
|---|---|---|
| Query latency p95 / p99 | system.query_log | Ranks query shapes by tail latency, then tests sort-key, projection or materialized-view fixes against a replay of the same shapes. |
| Parts per partition | system.parts | Alerts at half of parts_to_delay_insert, then fixes batching or partition granularity rather than raising limits. |
| Merge and mutation backlog | system.merges, system.mutations | Flags stuck mutations and merge pressure before they turn into disk exhaustion or replication lag. |
| Replication queue and delay | system.replicas, system.replication_queue | Tracks absolute delay per replica and queue size so failover targets are always current. |
| CPU hot spots | system.trace_log, EXPLAIN PIPELINE | Profiles where the pipeline spends time and whether parallelism is actually used. |
Streaming and CDC
Ingestion pipelines that keep ClickHouse fresh without drowning it
Most ClickHouse incidents we are called into start at ingestion, not at query time. Small, frequent inserts create too many parts, merges fall behind, and the cluster spends its CPU compacting instead of answering questions.
ChistaDATA Server standardises the path from source to table:
- Apache Kafka through the Kafka table engine or sink connectors, with consumer lag tracked per partition.
- Debezium CDC from PostgreSQL, MySQL and SQL Server into ReplacingMergeTree or CollapsingMergeTree designs that handle updates and deletes correctly.
- Kinesis, RabbitMQ and Apache Flink through buffered sinks sized to the target batch.
- S3 and GCS files through
s3Queueand table functions for continuous file ingestion. - Asynchronous inserts where clients cannot batch, so the server does the buffering.
Then materialized views pre-aggregate the hot paths, so dashboards read thousands of rows instead of billions.
Reliability engineering
High availability and disaster recovery with ChistaDATA Server
A replica you have never failed over to is a hope, not a plan. ChistaDATA Server treats availability as something you measure: replication delay, Keeper quorum health, backup age and the time a restore actually takes.
ClickHouse Keeper uses Raft, so quorum arithmetic is fixed: a 3-voter ensemble tolerates one failure and a 5-voter ensemble tolerates two. We place voters across failure domains, keep Keeper off the busiest data nodes, and never run an even number of voters.
For regional DR we combine asynchronous replicas in a second region with checksummed full and incremental backups in object storage, and we run restore and failover drills every quarter with the customer’s team watching.
Every table is created with an explicit engine definition. We never rely on defaults in customer estates, because defaults change between releases.
CREATE TABLE analytics.events ON CLUSTER 'prod_cluster'
(
event_date Date,
event_time DateTime64(3, 'UTC'),
tenant_id UInt32,
event_type LowCardinality(String),
user_id UInt64,
payload String CODEC(ZSTD(3))
)
ENGINE = ReplicatedMergeTree('/clickhouse/tables/{shard}/analytics/events', '{replica}')
PARTITION BY toYYYYMM(event_date)
ORDER BY (tenant_id, event_type, event_time)
TTL event_date + INTERVAL 90 DAY TO VOLUME 'cold'
SETTINGS index_granularity = 8192, storage_policy = 'hot_cold';| Failure scenario | ChistaDATA Server design | How recovery is proven |
|---|---|---|
| Single node loss | Two or more replicas per shard, load balancing away from unhealthy replicas | Replica kill test; replication delay back to baseline |
| Keeper voter loss | 3 or 5 voters across failure domains | Voter stop test; inserts continue, system.zookeeper_connection healthy |
| Zone or data-centre loss | Replicas spread across zones; DR replicas in a second region | Quarterly failover drill with measured RTO |
| Logical corruption or bad DDL | Point-in-time backups, retention matched to the audit window | Quarterly restore to an isolated cluster with row-count validation |
Security and compliance posture
Security and governance built into every cluster
ChistaDATA Server ships with access control configured, not bolted on. We use SQL-driven RBAC, row policies for multi-tenant isolation, quotas and settings profiles to stop runaway queries, and TLS on client and inter-server traffic.
Authentication integrates with LDAP and Kerberos where your estate requires it, and every login and query is auditable through system.session_log and system.query_log.
- Least-privilege roles for applications, analysts and operators
- Row and column restrictions for regulated or multi-tenant data
- Encryption in transit, and encrypted disks for data at rest
- Credentials held as secrets, never in configuration files
- Controls mapped to GDPR, HIPAA, SOX, PCI DSS and SOC 2 requirements, with your auditors owning the certification
Vendor-neutral view
ChistaDATA Server vs DIY ClickHouse vs a managed cloud service
We are vendor-neutral by principle, and we will tell you when a different model fits better. A fully managed ClickHouse service can be the right answer for small teams with no residency constraints. ChistaDATA Server is built for estates that need control of the infrastructure and a team accountable for the outcome.
| Dimension | DIY open-source ClickHouse | Fully managed cloud service | ChistaDATA Server |
|---|---|---|---|
| Where data lives | Your infrastructure | Provider’s infrastructure or tenancy | Your infrastructure or your cloud account |
| Engine | Open-source ClickHouse | Provider build; some features differ from open source | 100% open-source ClickHouse, upstream compatible |
| Upgrades | Your team | Provider schedule | LTS-pinned, risk-assessed, rolling, on your schedule |
| 24×7 expertise | Only if you hire it | Platform support; schema and query design usually yours | ClickHouse engineers on call, S1 in 15 minutes |
| Exit cost | None | Migration project | None; the cluster stays yours |
Reference patterns
Workload patterns ChistaDATA Server is engineered for
These are reference designs we use as starting points, not customer stories. Each one begins with a measured assessment of the actual data and queries.
Real-time fraud and risk analytics
Sub-second aggregation over transaction streams from Kafka, with ReplacingMergeTree for late updates and strict row-level access.
Impression and attribution analytics
High-cardinality event tables, materialized rollups per campaign and multi-region replicas for client-facing dashboards.
Logs, metrics and traces
Wide event tables with aggressive codecs, TTL to object storage and query quotas that protect the cluster during incidents.
Industrial time-series
Per-device sort keys, downsampling views and retention tiers that keep years of history at a fraction of hot-storage cost.
Engagement model
How a ChistaDATA Server engagement runs
Measure the estate
Query-log profile, part health, Keeper and replication review, capacity and cost baseline. Delivered as a written report. See our ClickHouse performance audit.
Fix the foundations
Schema and sort-key redesign, ingestion batching, HA/DR topology and security baseline, staged and reversible. See ClickHouse consulting.
Run it 24×7×365
Monitoring, on-call, upgrades, backups and drills under an SLA. See ClickHouse managed services and ClickHouse DBA.
Keep improving
Quarterly reviews of p99 latency, cost per query and growth, with a roadmap for the next two quarters. See Data SRE.
Moving from Redshift, Snowflake, BigQuery, Druid, Pinot, Vertica or Elasticsearch? Our ClickHouse migration practice plans schema translation, dual-running and cutover with a rollback path at every step.
Commitments
ClickHouse support tiers and response SLAs
ChistaDATA Server customers are covered by the same enterprise severity model as our ClickHouse support contracts. Response clocks run 24×7×365, including weekends and public holidays.
15 minInitial response
Production down or data at risk
Typical exampleKeeper quorum lost, inserts rejected cluster-wide
12 hoursInitial response
Severe degradation, workaround possible
Typical exampleOne shard lagging, p99 latency several times baseline
24 hoursInitial response
Non-critical defect or performance question
Typical exampleSlow report query, merge tuning request
48 hoursInitial response
Advice, planning or documentation
Typical exampleUpgrade planning, schema review
FAQ
ChistaDATA Server for ClickHouse: frequently asked questions
Is ChistaDATA Server a fork of ClickHouse?
ChistaDATA Server is 100% open source and fully compatible with upstream ClickHouse. Tables, SQL, client drivers and wire protocols behave as they do on stock ClickHouse, so you can move between ChistaDATA Server and community builds without data conversion.
Which ClickHouse versions do you support?
We support the current LTS and stable lines and plan upgrades against the upstream calendar of two LTS releases a year, each supported for a year. Release coverage: ChistaDATA delivers consulting, 24×7 support, managed services and remote DBA on ClickHouse 26.9 (latest stable release, September 2026), 26.8 LTS (current long-term-support release, August 2026, supported to August 2027) and 26.3 LTS (supported to March 2027), plus every supported stable release in between, with upgrade engineering for 24.x and 25.x estates, including 25.8 LTS, which reached end of support in August 2026. Last verified October 2026.
Can ChistaDATA Server run on Kubernetes?
Yes. We run ClickHouse on Kubernetes with an operator, on virtual machines and on bare metal, across AWS, Azure, GCP and on-premises data centres. The choice follows your platform standards, not ours.
Do you replace ClickHouse Keeper with ZooKeeper?
New builds use ClickHouse Keeper by default. We still support existing ZooKeeper estates and plan migrations from ZooKeeper to Keeper as a staged, reversible change.
What does it cost?
Pricing depends on cluster size, coverage hours and the scope of engineering. Most engagements start with a fixed-scope assessment. Contact us for a proposal.
Will we be locked in?
No. Your data stays in standard ClickHouse formats on infrastructure you control, and every runbook we write for your estate is yours to keep.
Next step
Put ChistaDATA Server under your ClickHouse estate
Tell us about your cluster, your latency targets and your worst week of the last quarter. A ClickHouse engineer will reply with a measured view of what to fix first.