Andy Pavlo joins ClickHouse to establish ClickHouse Labs (clickhouse.com)
212 points by nikolay_sivko 6 hours ago
a34729t 3 hours ago
I'm very curious about the convergence of the best in class fast OLAP products (StarRocks, ClickHouse) with Trino. It sounds like everybody is going for decoupled compute/storage, using S3 or similar as the storage layer, and thus forgoing colocated joins (ok I know ClickHouse joins suck)...
So what does this mean for ingestion (and indexing)? Iceberg V3? Paimon? Bespoke ingestion through the DB engine to do the indexing?
jimmyl02 2 hours ago
It's also interesting how Clickhouse / Starrocks can now also act as a query planner and executor on top of non-native formats (ex. Iceberg).
I assume the native formats will always be faster / more optimized but the need for Trino as a separate executor while running either of these databases seems to be close to gone.
a34729t an hour ago
The benefit of StarRocks/ClickHouse over Trino is that you get secondary indices, but that means you have to do the indexing somehow.
Native format is faster (especially for colocated joins), but it's way more expensive if you have to run a bunch of separate storage nodes vs just using S3, especially your query volume isn't that high.
I liken it to the BigQuery cost model, where storage is effectively free.
efromvt 2 hours ago
I’ve historically read this as ‘open format compatible’ but ‘native preferred’ - where this opens up market space and dev velocity - but it’ll be interesting to see if native storage differentiation gets dumped entirely. It just seems like ‘fork and optimize for our engine’ would always be tempting enough that you’d want a native play for when you don’t need the decoupling.
mrits 2 hours ago
Clickhouse joins have been improving almost every month for the last couple of years. Maybe they still suck but a lot less
remywang 2 hours ago
Looks like Andy's here, so if you see this - please also try to convince ClickHouse to consider funding DB research in academia. With all the money being poured into AI and the chaos in government funding, there is almost nothing for DB research anymore.
tomsanbear 5 hours ago
Always enjoyed his lecture series from CMU, hopefully those continue in a sponsored format from Clickhouse.
apavlo 4 hours ago
They will continue. New seminar series starts next month (announcement coming this week).
AbbeFaria 3 hours ago
I “audited” CMU 15-445 in April 2026, a few months ago. I then applied to Azure Hyperscale DB team, and as my pitch to get hired I cold DM’d the hiring manager whatever I finished in the assignments (BPM, Concurrent B-tree, Transaction support etc).
With that, I was able to get the interview although I ended up failing it. My lack of experience in C++ was probably one of the main reasons. Your course helped me stand my own during the interview and even though I had zero database experience apart from finishing the course, I felt adequately prepared.
Thank you for making the course open source. As a working professional, it was helpful to be able to do the course on my own time. Suffice to say, I am a big fan of your work and thank you for introducing me (and other fellow learners) to the interesting world of Databases!
ibgeek 2 hours ago
Will you maintain a connection to CMU? If so, what will you continue doing and at what percent of effort?
coolThingsFirst 27 minutes ago
Your DB course is phenomenal. Thank you very much for making it available to all of us!
danlark1 5 hours ago
Congrats, that's incredible news. Watched apavlo@'s lectures while studying in the university and finished my bachelor thesis implementing features and doing research at ClickHouse. Surprised to see these worlds being together now!
gavinray 4 hours ago
Clickhouse just became the hottest talent-attraction on the market.
Congrats Andy, hope you enjoy the ride =)
adrianco 5 hours ago
Hi Andy, that’s very cool news. Will you be at the HPTS workshop in October?
throwaw12 5 hours ago
That's great news for Clickhouse and for Andy I guess, they are both obsessed with databases and are at the cutting edge.
Tostino 2 hours ago
Hope it goes well for you Andy. Have loved your lectures over the years.
Best of luck!
sigbottle 2 hours ago
Yooo is this the CMU funny guy?
nikhilsimha 4 hours ago
Congrats! Andy is incredible!
pphysch 5 hours ago
> The goal of ClickHouse Labs is to establish a best-in-class industry research organization focused on databases. It will not operate as an isolated research organization that throws ideas over the wall to engineering. Instead, we will work closely with ClickHouse engineers, customers, collaborators, and industry partners to develop and disseminate new ideas that keep ClickHouse at the bleeding edge.
This is cool, though bittersweet that the public research infrastructure (universities) is not really configured to support this kind of high-impact research any more.
Ar-Curunir 2 hours ago
What do you mean? There’s plenty of high impact research on databases happening in academia. (Including, for example, Andy’s group at CMU)
Also keep in mind that the part you quoted is partially marketing copy.
bootwoot 2 hours ago
I take this announcement to mean he is leaving. So he's a perfect example
kenferry 26 minutes ago
sghiassy 5 hours ago
Clickhouse Labs is a research lab for ClickHouse’s engineers and their db product?
Whatever floats your boat. Sounds like you just work as an engineer at a db company
lumost 5 hours ago
Deep tech fields have a more complex relationship with academia than other industries. Engineers often need license to test ideas that may not see adoption for years, or require tens of millions to build-out.
A "research" arm is exactly this license, although it comes at the cost of potentially killing innovation in the rest of the company.
ForHackernews 5 hours ago
C'mon, databases are great but they are not in any way a "deep tech" field. Databases exist today in a zillion production forms as commercial products and free software.
Deep tech includes things like nuclear fusion, solid state batteries, quantum computers. I know everyone wants to feel cool, but just because your new javascript framework will be in beta for the next ten years doesn't make it "deep tech".
jandrewrogers 4 hours ago
mamcx an hour ago
lumost 5 hours ago
apavlo 5 hours ago
I'm not sure what you think "research lab" means? As I mentioned in the article, we look at IBM Research (Almaden) and Microsoft Research as inspiration.
ForHackernews 5 hours ago
Looking forward to the database analogues of IBM Watson and Tay.ai
I'm partly kidding; there's plenty of space for innovation in databases. You should sponsor https://sled.rs/
tclancy 3 hours ago
jandrewrogers 5 hours ago
Many database companies have research labs. It allows you to explore ideas, unorthodox concepts, and parts of the design space that may never be reflected in an actual product. Or to figure out how to solve specific hard problems that you come across with neither a good solution nor proof of impossibility in literature.
This is essential for a company that wants to stay on the frontier of database tech.
bedman12345 4 hours ago
Which database companies have research labs? AFAIK amazon, snowflake, databricks, google and SAP don’t have a research lab dedicated to databases. They have some people that are paid to do research but there does not seem to be an IBM Almaden anywhere in the world.
senderista an hour ago
astrange 5 hours ago
> ClickHouse had features that at the time were only found in a handful of closed-source, commercial analytical DBMSs. For example, ClickHouse was written in C++ and supported vectorized query execution using SIMD in 2016. Most prominent open-source analytical DBMSs in 2016 were JVM-based and did not support SIMD optimizations until years later.
Performance is a feature. "Written in C++" is a strange idea of a feature.
bijowo1676 4 hours ago
zero cost abstraction is one of the many features of C++, its just you are not familiar with the language