I've operated petabyte-scale ClickHouse clusters for 5 years (tinybird.co)
182 points by adastral 4 days ago
zbentley 11 hours ago
> For loads with over 20k rows/s and people pushing changes, you may need a full-time person to handle the cluster and take a look at the crazy queries people are going to write.
I think this was a benefit of DBA culture in previous eras. Not that the DBAs were specifically necessary to write good queries (often they'd need to work with application teams to guide them towards schemas/behavior that worked well) or to maintain the database (managed DB offerings obsolete a lot of this work), but because they functioned as gatekeepers and rate-limiters of what queries and schemas could exist.
In that mode, DBAs functioned a bit like a human/process version of a thin microservice wrapping database access functionality. A big benefit was that the rate of change of queries/schema changes/access patterns was controlled and had a higher probability of being reviewed and thought about by humans before it went live. This also resulted in an increased end-database-user culture of trying to make existing schemas/query patterns work before jumping straight to bespoke access patterns. That culture's not what you want as e.g. a startup or pro-rapid-big-refactors shop, but it is what you want when your DB reliability needs or query rate/dataset size are high.
I don't think it's a given that a gatekeeper team is worth the overhead and cost; that's situational. I do think that the code version of that team (aforementioned microservice that wraps DB accesses/schema changes and nothing else) is usually not worth the cost. In my experience, that pretty much always reduces reliability and free performance gains that come from using direct DB clients from user code.
bushbaba 11 hours ago
Most startups can just scale your traditional separation of compute & storage here though. You’d be shocked how well duckdb against s3 scales for 99.9% of use cases
zbentley 11 hours ago
It's one thing to be able to scale database compute/storage; it's another thing to be able to partition it. It's extremely common for bad queries/access patterns to cause noisy-neighbor effects on other simultaneous accesses to the database, to the extreme of knocking the whole database over with timeouts/OOMs/etc.
Scaling out DB compute can only help with that to a (expensive) point; eventually, you end up wanting to either prevent the bad queries from being added to the system (DBA culture) or ensure that the bad query runs on database infrastructure that doesn't affect other queries. That's why partitioning DB compute (and storage: noisy-neighbor effects from a bad query at the storage layer don't require storage to be running e.g. a BookKeeper or whatever on a server; they can manifest as hot S3 keys or cloud object/block store rate limiting) is a necessary capability if your plan for dealing with a culture of "anyone can add any access pattern they want" is to scale the DB.
bushbaba 10 hours ago
jeremyjh 10 hours ago
bigfatkitten 7 hours ago
Most software developers now are absolutely ignorant of performance concerns. Just throw more compute at it until it works, and someone else will pay the AWS bill.
0xbadcafebee 11 minutes ago
winrid 18 minutes ago
hello we still exist :)
tmpz22 11 hours ago
Instead we're... listen to this... we're going to take a software developer right. Just a normal developer right. We're going to make them be the database expert right. And the cloud expert. And we're going to put them on call. We're going to have them debug linux logs, and optimize our AWS costs. They'll be there for client escalation work. And big sales calls. From time to time we'll even have them do front end work.
And get this. We pay them the exact same.
zbentley 11 hours ago
I think a DBA/ops/infrastructure person as an imposed bottleneck is a useful capability in some environments.
But I won't follow you as far as "expecting developers to have expertise in how and where their software runs is unreasonable".
Like, yeah, it sucks that added DevOps responsibilities etc. don't come with adjusted compensation/time allocation expectations. I'm with you there.
But it's simultaneously true that a ton of "just regular developer" people are significant liabilities because they don't understand anything about the environment where their software runs. That liability manifests operationally (if someone's just running integration tests on Windows for their Java business logic changes and don't have any familiarity with e.g. the Linux, container, or cloud environments where their code runs, they're going to be useless when their code breaks in production and operations staff needs context), and it also makes them less effective when writing code--this culture of "developers should just live in business logic and not have to context-switch or fill their brains with other levels of the stack" is what leads to full table scans, lack of awareness of memory use, N+1 query hell, looping microservice dependencies, misunderstanding of what HTTP fields are set on requests that are mutated by load balancers, mistaken assumptions about how many instances of code can run and what concurrency/thread/coroutine behaviors are present, and so on. Those are very common problems, and it's incumbent on developers in every specialty to gain familiarity with how and where their code runs in order to write and maintain that code effectively.
If your code runs on Linux in Kubernetes, all of your developers should know how to read Linux system logs, check database sessions/queries issued by parts of the application, ls/grep/cat/strace/ps their way around, interpret k8s/application dashboards, check application logs both in log storage and as they're emitted from a process, exec into a container, restart pods, check deployment liveness, etc. Even if they don't have permission to do those things in production.
That was true in 2005 when they deployed their code to IIS on Windows Server/MSSQL, too--just with different operational specifics.
That's a low bar that's often unmet, and all sorts of teams suffer from that failure. Those skills can be trained, kept up to date, and hired for; I don't think there's a great excuse for not expecting them.
kentm 9 hours ago
tmpz22 11 hours ago
mawadev 9 hours ago
Eerily accurate how it works these days, I wish you weren't correct. I met a DBA wizard (he looked like the creator of c++) at a banking IT dept and this guy intuitively sensed what you needed and how its done.
Foobar8568 9 hours ago
And don't forget contract management with the supplier, L1-L2-L3 support, all in one, and integrating as the supplier is useless and your contract is shit.
Oh and you will do also business analysis with the business as claude is too complex for them (read any version of the nocode initiative).
FLeXMurphy 10 hours ago
I read this in Steve Jobs voice. But maybe I was supposed to use Dr. Ian Malcolm instead?
neya 9 hours ago
> And get this. We pay them the exact same.
Why are you complaining? You should be grateful of the learning opportunity not everyone gets to have. Sure, we pay you peanuts for it. But, it's ultimately for your own good. Don't mind my yacht and Ferrari, though.
PunchyHamster 6 hours ago
That kinda how moving from programmers + sysadmins to devops looks like.
Managers went happy coz now they don't need to have hire sysadmins, while in reality they hire sysadmins, call them devops, and have them know some programming on the side.
And the "savings" from not having onprem infrastructure are burned on expensive cloud and debugging cloud blackboxes
throwaway894345 11 hours ago
A decade ago we would hire them fresh from some Ruby on Rails bootcamp so we could pay them less :shrug:
chasd00 11 hours ago
yes, that was the cloud and "devops" promise. ..or what it just another sham?
WarcrimeActual 10 hours ago
>Instead we're... listen to this... we're going to take a software developer right. Just a normal developer right.
Maybe I'm old, and I am, but I just can't get past this point with such annoying writing. Like if you actually spoke like this people would hate you.
dwedge 10 hours ago
ezekiel68 7 hours ago
pstuart an hour ago
I'm going to go out on a limb here and say that having a tuned LLM would probably be able to eliminate the need for a full time DBA query analyst.
Disclaimer: I've never managed a ClickHouse cluster, let alone one of this size.
walthamstow 12 hours ago
As an aside, I was stuck when turning on the Fulham v Crystal Palace game last week to find that Fulham have ClickHouse on their shirts this year, and Palace have Temporal AI. Talk about my worlds colliding.
lucrbvi 12 hours ago
That's a lot of ®, curious how ClickHouse® Inc. is treating the use of its name by others ... Hopes it's not like Oracle with JavaScript
HatchedLake721 11 hours ago
They sell managed ClickHouse so I suspect it’s a precaution
gcharbonnier 8 hours ago
From the article: > Let me clarify something: we created Tinybird ClickHouse because we wanted to build an analytics application without all the pain I'm describing. We do not offer ClickHouse® hosting; we solve the analytics problem
EDIT: I'm wrong, their home page clearly states: "Ship fast over a Managed ClickHouse®", though I don't really understand the difference between a managed clickhouse and clickhouse hosting...
doe88 11 hours ago
It's defensive language for sure, i don't know how much it adds of protection in reality, but i nonethelesss sympathize with the author if he feels the need to protect himself that way or signaling the risk he takes.
mritchie712 7 hours ago
> Tinybird is not affiliated with, associated with, or sponsored by ClickHouse, Inc. ClickHouse® is a registered trademark of ClickHouse, Inc.
Yeah, it sucks they need to do this. If I was a visitor to their website, I'd immediately want to know what ClickHouse, Inc. is and you'd realize ---> it's managed clickhouse, direct competitor... why would I use the one that needs all the ®'s
wackget an hour ago
What kind of systems are generating enough useful analytics data to ever scale to petabytes in the first place?
It's difficult (for my bird brain, at least) to imagine a scenario where such volume of analytics data would ever be necessary.
x0x0 31 minutes ago
It's multiple records per user over time. See something like posthog. One interaction with a form can generate 10-15 events. eg enter page, click link with text "settings", change input, clicked input, etc. So active users can generate 15-150 events per session, each ~1kb to 2kb in size.
The reason to store that data is to be able to build analytics and funnels you didn't anticipate and construct them retrospectively. You can of course save massive amounts of storage by building rollups and aggregating, but that prevents you from being able to eg change the funnels and see the funnel in the past.
10M users × 100 events/day × 1KB ≈ 1TB/day. A couple years gets you into PB range.
The other obvious thing is logs. Being able to have a vulnerability then go back in time and see if you were exploited is super valuable. See eg log4shell: if you kept records, you probably could go back and see if you'd been exploited.
bradleyy 12 hours ago
I just wish Amazon would offer it as an RDS DB; it'd make my life so much easier.
andriy_koval 11 hours ago
You can use actual CH Cloud on AWS?..
bradleyy 11 hours ago
I'm aware; unfortunately that then leads to "must have a vendor approval" and a lot more process. If it were RDS, then it'd just be provision and done.
bijowo1676 5 hours ago
sdairs 7 hours ago
fidotron 11 hours ago
Do any two teams actually operate it in anything like the same way though?
What I saw of it, especially some years ago, was it was highly particular, and everyone had their own odd habits built around running it, ingestion, querying, everything, to the point I suspect there are a non trivial number of companies using it where it is actually the core operational expertise of the company, despite them all appearing to be in totally different domains.
f311a 11 hours ago
You can use ClickHouse cloud to host it on AWS.
But given the majority of use-cases of CH, AWS can be quite expensive.
hodgesrm 10 hours ago
Using BYOC management reduces the costs significantly. The big cost in analytic SaaS offerings is generally compute, which vendors mark up significantly. (They keep margins low on storage.)
Disclosure: My company Altinity offers BYOC management of ClickHouse.
threecheese 11 hours ago
I’m operating a terabyte-scale Clickhouse - but only because I left Langfuse running for a few months on a MacBook :)
But seriously Clickhouse does love disk space.
cnkk 9 hours ago
But on the same time it is super efficient with it compare to other solutions. Kind of efficient for application logs for example.
anguss 8 hours ago
A couple weeks ago I was scanning Github for repos with frequent commits to identify so-called "software factories" and was surprised to see clickhouse. I wouldn't touch it with a 10ft pole considering how quickly they're merging code into main. I'm talking 50+ commits per day and thousands of AI generated issues and triages. Check it out for yourself https://github.com/clickhouse/clickhouse
ryan_lane 4 hours ago
Just wait till you find out how many commits per day to a product are occurring to most of the SaaS you rely on.
Clickhouse has a managed SaaS and it's their primary product. They have a lot of engineers. They're going to do a lot of commits.
yakkomajuri 11 hours ago
> "Every single company handling ClickHouse® struggles with ingestion."
Very true. Reading about "too many parts" gave me flashbacks.
(previously owned ingestion into CH at PostHog, no longer)
fuziontech 9 hours ago
Scaling during that era was fun. Slack was non-stop:
:oof-1: CH needs more disk to keep up with merges.
We ran CH way too lean in those days.
yakkomajuri 3 hours ago
ofc we meet in the ClickHouse thread
pjot 3 hours ago
I don’t think it’s necessarily the use case changed but maybe the constraints. Streaming directly to Clickhouse is a great solution. Recent approaches I’ve seen though buffer data as parquet in object storage first, and read data from there. Im not sure if CH would be the preference
wackget an hour ago
Yeah that's cool but your blog won't load on my hardened Firefox:
- Uncaught TypeError: navigator.sendBeacon is not a function
- Uncaught (in promise) ReferenceError: WebAssembly is not defined
Not sure why a blog entry needs WebAssembly but I'll give it a pass.
phoghed 33 minutes ago
Sometimes I think you guys run your browser like this just to make these performative comments
dev_l1x_be 7 hours ago
Complex systems have complex operations and problems. I think isolating the read and write workloads and separating the read queries by departments or even users is a distinct possibility especially with technologies like DuckDB.
a34729t 4 hours ago
That is exactly where all the high scale stuff like Trino and StarRocks have gone (and ClickHouse can use S3 as storage too). However, the problem you run into is when you want to index stuff- secondary indices are tightly coupled to the query engine, so in practice you arent going to use StarRocks to write and index data to the object store, and then use ClickHouse or Trino to query it. I think it would be a useful development to decouple writes and reads, but this requires a common indexing scheme, and then once you have this standard, in principle lots more indexing plugins can be written and all compatible query engines can use them. Apache DataFusion is probably the right framework to build off of.
mrngm 11 hours ago
For those interested, there's also a second part [2025]: https://www.tinybird.co/blog/what-i-learned-operating-clickh...
(note: the first part was originally published April 2025 according to the date tooltip)
Perepiska 5 hours ago
Some entity spoiled the link "they don't like it" by adding (R), correct link is [1]
1. https://github.com/ClickHouse/ClickHouse/issues/49383#issuec...
Lucasoato 11 hours ago
Where does ClickHouse fit between ElasticSearch, Pinot, TrinoDB, or just plain Spark? These are very different tools but I’m curious to know if any one has already compared them and can share some thoughts regarding their maintainability, QPS, latency, etc...
a34729t 2 hours ago
You have to make a distinction betweet query engines, datastores and batch compute.
Trino is a federated query engine and lacks secondary indices, but can do all sorts of big scale stuff like distributed merge sort and spool to disk. The typical use case is querying your data lake (say Iceberg or Hive catalog with a bunch of Parquet format files in S3, and a separate ingestion system). It is extremely mature and well understood with lots of extensions.
ClickHouse, Pinot, StarRocks, and Druid are full on OLAP datastores that can handle varying rates of ingestion (Druid is by far the slowest, the rest are fast like 100kqps writes per table is nothing fast). ClickHouse is the most used but sucks at joins and stateful data. Druid, Pinot and StarRocks can do joins and handle stateful data. In my experience Druid is the worst, Pinot is relatively immature and has minimal industry traction, and StarRocks is the most mature and has lots of traction, both in China and in ethnic Chinese analytics teams in US companies. They can all support high qps for trivial queries (thousands of simple queries in tens of ms, given enough hardware), but depending on data partitioning it can get slow quick handling a bunch of concurrent queries that are scanning the same physical servers. But people have PostGres vertically scaled to 100kqps plus and 100+TB too. In principle the use case is dashboards and charts for your real time UI; tier down to S3 with Trino for more flexible, bigger queries.
The cool thing about all these systems (the fast OLAP) systems is that they are all converging with Trino where they are moving their storage layer to object storage, which is way more flexible. No more hot storage nodes, and infinite storage. And then use Paimon or Iceberg v3 as your data lake and you get near real time stateful storage.
ElasticSearch is amazing but is really optimized for full text search and aggregations, and while it can scale to huge sizes, it does not give you joins and forces you into a very particular approach to materialized views. Also, not columnar... and nested documented dont scale well in my experience.
Spark on the other hand is just your old batch compute framework at this point. It is very flexible, and you are writing a series of SQLish transformations, but for many simpler use cases Trino is much faster and easier to use. Way bigger learning curve than just writing SQL and hoping your query engine has a good optimizer!
hodgesrm 10 hours ago
The relationship depends on the use case and is not linear. For web analytics and log management ClickHouse is a great replacement for ElasticSearch, for example.
antoniojtorres 11 hours ago
I would put it closest to Pinot though pinot makes different architectural choices about storage as a difference that stands out to me. Then i’d rank trino as a second closest in only some ways. Spark the furthest for sure.
AtNightWeCode 9 hours ago
ClickHouse I would say is more for analytics and data warehouse types of loads. So not a direct competitor to those tools except maybe for Pinot.
It is very easy to ingest data into CH. We connected exchange topics and it just worked with zero code.
But the thing with CH is that it is pretty much a Russian product so you should not use it for production anymore.
AntonFriberg 9 hours ago
I would like to see some sources on your claim about it being Russian. It is incorporated in the US with most developers in Amsterdam.
After the invasion of Ukraine they stayed silent for a while but that was because they needed to allow there developers to get out of Russia. Many of them are Ukrainians including the CTO and founder. As soon as it was safe for there team they took a very firm stance against Russia with Ukrainian flags on the website and written statement from the team.
I am not aware of any Russian influence currently.
AtNightWeCode 8 hours ago
Lucasoato 8 hours ago
We’ve come at a time in which we fear exploring projects like Apache Doris because we fear that OSS might be used in malicious ways by state-backed actors.
trynotsober 11 hours ago
When replaying customer queries against the next version, how do you compare results for queries using now() or approximate aggregates? Curious how you separate expected differences from actual regressions.
solatic 9 hours ago
I cocked an eyebrow more than once reading this.
> A quick note about HTTP: ClickHouse® offers a TCP connector with a native protocol, but we don't use it. It does not offer many advantages for the type of application we build
This needs more elaboration. One of the major goals of running a ClickHouse cluster is to provide low-latency queries; a persistent TCP connection removes the need to re-establish a new connection for each query and thus reduces overall latency in line with CH goals. So I really didn't understand this.
> ClickHouse® open source faces a significant challenge: limited support for cloud storage. Modern OLAP databases and data systems should leverage cloud storage for cost efficiency and independent scaling of compute and storage resources. Snowflake established this standard over a decade ago, and ClickHouse® (open source) lags behind
ClickHouse writing to NVMEs is exactly how they provide their latency and performance advtanges. Writing and reading to object buckets is fundamentally slower with multiple network hops to reach what is, in this architectural context, a storage server for your storage server. If you really need far more storage, and are willing to sacrifice query latency to get it... why not architect for one of the OLAP databases, like Snowflake, where that was part of their architecture from day one?
> Because you are testing your analytics queries, right?
No? Half the point of an OLAP database is to let users write their own queries. If we knew the queries ahead of time, we probably wouldn't need an OLAP database, and instead use a less-flexible streaming architecture storing intermediate calculations so as not to need to pay for petabyte-scale storage. The expected value from paying for all of that storage is to support not knowing which queries will be written by users.
> Every single company handling ClickHouse® struggles with ingestion... Backpressure mechanism: Some people put Kafka before ClickHouse®. This does the job
The whole trade-off that you make with column-store databases like ClickHouse (instead of row-store databases like Postgres) is that inserts are slow for column stores (whereas they are fast for row stores). Inserts happen slowly, asynchronously, in the background. It is the price you pay for fast analytics queries. This is why OLAP databases have a latency lag and do not show real-time results. This is why stores like Kafka are usually a good fit, you let Kafka hold onto new data until batch insertions can catch up. If you do need real-time queries, you don't write to an OLAP directly; you write to a stateful frontend that answers the query itself, then streams out historical data from the OLAP that was successfully written there. And the first thing you do in a "I want to have my cake, and eat it too, and yes I'm willing to pay for the privilege" architecture like that is... to keep the persistent TCP connections, because that's really low-hanging fruit.
lee_ars 7 hours ago
My dude, you only need to put the registered trademark symbol after the first occurrence of the mark. You® don't® need® to® put® it® after® every® mention® of® Click®House®™™®™®™®™®™.
will_pseudonym 10 hours ago
At first when I read this, I was like "Why does The Onion's ClickHole site need so many servers?"
alexnewman 5 hours ago
About 20 years ago the titans of the database fields told us they wanted to put ML in the brain of databases. I think agentic databases are the future.