AI Expands the Role of Data Scientists and Strengthens Governance

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Code automation, agents integrated into workflows, accelerated analyses: AI is no longer just saving data scientists time; it is redefining their scope of action. Meanwhile, data governance, management of agent skills, and the redesign of BI tools are becoming central projects.
Governance of Semantic Layers and Agents Becomes Ongoing Work
The widespread adoption of AI makes data governance more critical. Semantic layers must be maintained, as definitions, structures, and context evolve over time. An agent can periodically check the documented structure and aggregate new knowledge from internal tools like GitHub, Google Docs, or Slack, but human validation remains necessary. The author notes having participated in building the semantic layer at their previous company and is considering best practices for starting from scratch as the first data employee at a new startup. Agent workflows also require governance, as skills can easily multiply: they cite a team repository that grew to over 100 skills in three months. Reusability demands high quality and reproducibility, with reading and testing before sharing. An agent can also flag redundancies and clean up the context.
Redesign of BI and Proliferation of Dashboards to be Managed
Business intelligence is undergoing a redesign phase. Reference reports and drag-and-drop exploration are being questioned, with AI offering more flexible ways to create and share dashboards through tools like Claude Artifacts, Databricks applications, or internal HTML pages. The value of traditional BI tools is increasingly being challenged, even as some add AI agents or reposition themselves towards agent-assisted exploration, which, according to the author, does not resolve the underlying crisis. At the same time, the ease of producing dashboards complicates the management of the "source of truth": similar versions multiply, sometimes without incorporated underlying code, complicating validation and maintenance. Discrepancies in figures may arise from legitimate filters rather than errors. A strategy of verified dashboards by domain remains valid, and it is recommended to include or deposit the code to allow for human or agent reviews.
Faster Analyses Through AI-Assisted Simulations and Causal Inference
AI significantly reduces analysis times, as illustrated by the example of geographic tests. Previously, it was common to use a Databricks notebook, manually set metrics, market lists, and thresholds, with possible errors, and then interpret results through DiD and regressions. Now, they entrust the model code to AI, describe their constraints in natural language, and let AI execute simulations that select treatment and control markets. AI is also deemed useful for comparing different causal inference approaches to strengthen the robustness of measurements.
AI Invites Itself to Every Step of the Analytical Workflow, from Framing to Reporting
AI is no longer confined to code generation. MCPs connect tools to gather and summarize the history of discussions and research. A planning mode is used to debate methods and establish the analysis plan. AI then executes the plan, highlights intermediate results, and produces a final report for stakeholders.
Expanded Scope Towards Engineering Thanks to AI Tools
AI enhances data scientists' ability to take on engineering tasks. Six years ago, in a startup, it was common to combine engineering, analysis, and data science functions, following a checklist covering dbt, YAML, Docker, test commands, triggering Airflow, and monitoring via Datadog, with a time-consuming environment setup. In their current role as the first data employee, they indicate that determining the data model and logic is sufficient, with Claude Code managing implementation and deployment with unit tests. They report similar benefits for developing machine learning models and other tasks, made quick with suitable tools and code reviews. According to them, AI does not just "accelerate": it expands what a data scientist can legitimately take on.
Stakeholder Self-Service and Increased Role of Semantic Layers
A movement towards self-service is emerging: the author claims to have coded very little manually in six months, with AI generating most of the analysis code while their role becomes that of a reviewer. Repetitive workflows are encapsulated in agent skills, from analyzing causes of metric variations to drafting experiment reports by connecting to an MCP tool. Stakeholders, particularly technical ones, are taking charge of tools like Claude Code to extract figures and conduct exploratory analyses when time is short on the data side. This autonomy raises questions about reliability, hence the growing importance of semantic layers. The author states that these layers, by documenting tables, metrics, and dimensions, help agents extract indicators like a retention rate without piecing together scattered definitions, which, according to them, improves accuracy and saves tokens. They observe divergent reactions to AI—loss of trust after errors or excessive trust leading to incorrect figures being cited—and argue that structured context is key. The bottleneck shifts from data access to the trust placed in the results.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.