Data Analysis & Visualization at Scale on Semi-structured Data
seen from Germany
seen from Singapore

seen from United States

seen from Libya

seen from United States
seen from United Kingdom
seen from United States

seen from Germany
seen from Jordan

seen from Brazil
seen from United States
seen from United States
seen from United States
seen from United States
seen from Germany
seen from United States

seen from United States
seen from United States
seen from United States

seen from Italy
Data Analysis & Visualization at Scale on Semi-structured Data
When Gabriel Straub first introduced me to Eric Ries’ book titled “The Lean Start-up”, I instantly knew that it should be applicable to the world of Data Science. However, translating the approach from the domain of a start-up business to that of a Data Science team or project was by no means obviou
Shipping GenAI isn't just models, it's 𝐫𝐞𝐩𝐞𝐚𝐭𝐚𝐛𝐥𝐞 𝐩𝐢𝐩𝐞𝐥𝐢𝐧𝐞𝐬 with 𝐠𝐨𝐯𝐞𝐫𝐧𝐚𝐛𝐥𝐞 𝐝𝐚𝐭𝐚 𝐩𝐚𝐭𝐡𝐬.
So 𝘧𝘳𝘰𝘮 𝘵𝘶𝘵𝘰𝘳𝘪𝘢𝘭 𝘵𝘰 𝘵𝘦𝘢𝘮 𝘱𝘭𝘢𝘺𝘣𝘰𝘰𝘬: I've distilled Databricks' unstructured data pipeline for RAG into a 𝒎𝒊𝒏𝒊𝒎𝒂𝒍, 𝒓𝒖𝒏𝒏𝒂𝒃𝒍𝒆 𝒘𝒂𝒍𝒌𝒕𝒉𝒓𝒐𝒖𝒈𝒉 your teams can drop into a UC-governed project and start measuring.
Highlights:
- widgets and a _bootstrap setup (that software developers might hate but data scientists will appreciate).
- gotchas: path configs and UC table naming.
- levers to tune: chunk size/overlap, model swaps.
Leader's angle: 𝐬𝐭𝐚𝐧𝐝𝐚𝐫𝐝𝐢𝐬𝐞 𝐞𝐱𝐩𝐞𝐫𝐢𝐦𝐞𝐧𝐭𝐬, reduce setup drag, and enable fast iteration and comparison of results across projects and teams.
This is a walkthrough of the Databricks tutorials for setting up an unstructured data pipeline for RAG (retrieval augmented generation)…