Updated
Updated · O'Reilly Media · Sep 30
Data Scientists Get 5-Step Playbook for Directing AI Agents
Updated
Updated · O'Reilly Media · Sep 30

Data Scientists Get 5-Step Playbook for Directing AI Agents

3 articles · Updated · O'Reilly Media · Sep 30

Summary

  • Five practices anchor the playbook: frame the investigation, equip the agent, organize experiments, review results independently, and turn lessons into reusable skills and evals.
  • A fraud-detection test showed why that structure matters: Claude first posted F1 of 0.87 and ROC AUC of 0.99 using a random split and a planted leaking feature, overstating real-world performance.
  • After the team enforced a temporal holdout, removed leakage and required subgroup reporting, F1 fell to 0.70, overall recall to 0.61 and high-degree-node recall to 0.21.
  • The article argues data scientists should shift from manually executing every step to specifying the question, runtime and constraints, then verifying evidence through bounded experiments or competing analyses.
  • That workflow extends to reusable infrastructure—skills, access controls, evidence trails and independent review—especially when colleagues use agents directly rather than through a human analyst.

Insights

As AI agents take over data science tasks, is the real job now designing trustworthy tests rather than building models?
If an AI agent can build a fraud model in minutes, who catches the hidden leakage before fake metrics become real decisions?
Why did a Bitcoin fraud model’s F1 score fall from 0.87 to 0.70 once humans enforced real-world rules the AI missed?