Writing the Data News — metaphor (credits)

Hey, I hope you're doing good. Let's jump to the Data News, this week is full of great content.

Data Fundraising 💰

Snowflake Summit

Last week I forgot to write specifically about the Snowflake Summit. It was in Vegas and people were excited like if Steve Jobs appeared in the warehouse. Some stuff has been announced there:

Snowflake founders (credits)

Data teams organisation — apply the JTBD framework

It's been a long time since I've not share thoughts on data teams. This week Emilie proposed to use the JTBD framework to build more effective data teams. The Jobs to be done framework is a way to prioritize work. As a data team our mission is to empowers people in their decision-making.

If you identifies correctly your Jobs, the team will obviously drive enablement on others to drive business impact. Emilie shared 5 frequents jobs she observed. I recommend you to read them all.

In some extend I also recommend you to read Christophe's post on Airbyte blog on how to structure a data team to climb the AI pyramid of needs.

Our journey towards an open data platform

Doron shares how they build their data platform at Yotpo. All the drawings are super useful and clear. When you look deeply at the platform you could have interrogations about the needs to have ~3 data storage (Redshift, Snowflake and Databricks) and 3 visualisations tools.

This is a good feedback post but it shows super well the technologies explosion — cf. state of data engineering — we face today in the data ecosystem. Tools cherry-picking is becoming an art.

A framework for designing document processing solutions

Data extraction from document is slightly becoming for a lot of companies one of the best way to apply artificial intelligence for the business and to help the operational teams. Humans love paperwork and all of this paperwork is just in demand to be parsed.

Lester James proposed a framework for designing document processing solutions. Converting PDF to usable data is a key task. To do that he proposed 3 steps: annotation, multimodal models and an evaluation step. As a disclaimer he showcases an annotation lib (Prodigy) he works on.

From Jupyter to Kubernetes: Refactoring and Deploying Notebooks

I imagine that Netflix energy spent to put in production Notebooks didn't stay unnoticed. People are developing Ploomber to help you doing data pipelines from Notebooks.

On the other side you can also try to measure the CO2 impact of your notebooks (on Azure).

Data scientists (credits)

Side projects FTW

I've always been a huge fan of side projects in order to learn something, so each time I see people doing extra stuff with data I take care to share it because it resonates in me.

This week Jack built a data pipeline for his own Strava data to visualise everything in Tableau. Almost 7200 kms in 2021, well done.

Fast News ⚡️

Podcast 🎙

A discussion about dbt, Airflow and the semantic layer. Good one.

Last Read — 3 years after Data mesh: lessons learnt

A month old article. Michelin detailed lessons learnt from implementing a Data Mesh — which, as a side note became suddently un-trendy this year. The main blocs are a data fabric [the data platform], distributed data domains w/ teams [exposing and/or storing and/or valuing data] and a federated governance.


See you next week. I Love you all.