Data engineering team moving data manually (credits)

Dear readers, I hope you had a great week. Each time I look back and I see the amount of Fridays I've spent reading and writing I'm still surprised. For the last 2 newsletters I've tried to ask your for paying support. From number of people who really paid I can see that I failed to either word it correctly, either to propose a newsletter where you see the value of paying for it.

In any case, I need your honest feedback, what would make you consider paying for the content I create?

This is something I struggle with, I really like writing, I really like this newsletter, I really like the blog, but it takes me one day per week to be done. If I want to continue for years I have to find a way to make it sustainable for me, and also if I want to continue more in this direction I have to find a model that works. I'm open to all honest feedbacks.

A bit of infrastructure

This week I've seen a lot of articles that I can put under the infrastructure category, so here we are. The current data state is heavily dependent on infrastructure, wether it's cloud, on-premise or semi-related we need to understand where the data lands and where the code runs.

First Bucky gave his thoughts about the state of infra in 2023. In a nutshell, Javascript is the future of everything, we say it for years, you write once and you run it everywhere—in the browser, on servers—then workflows systems are a key piece of every software architecture, we have a fragmentation of tooling and we want to run tasks one after the other which means we need something to orchestrate them, finally the OLAP databases are evolving in something different with many more features.

In order to improve your data infra you should sometimes try to occasionally kill your data stack, chaos engineering is something that helps discover issues. Monte Carlo also wrote this week about chaos engineering, with a manifesto.

When it comes to data storage, the real-time ecosystem has also changed a lot in the last few years and a lot of tooling went out to simplify the burden of managing Kafka clusters, Materialize—a real-time platform—detailed their architecture. But if you want to continue using the underlying tools here an overlook of Flink architecture or a few techniques you should know as a Kafka streams developer.

Finally Whatnot shared how the migrated their CD processes to ArgoCD and Pinterest now uses HTTP/3, which I didn't even know it was existing.

Is it Kafka? (credits)

Fast News ⚡️

Data Economy 💰


See you next week ❤️ — small edition today, blank page issues 🫠