How to build agents when the smartest AI isn't smart enough
Skills:
Agent Foundations80%
Key Takeaways
Discusses building agents when the smartest AI isn't smart enough, featuring Benchling AI's agent-based intelligence layer
Original Description
Nick Larus-Stone is the Head of AI at Benchling, the R&D data platform that life science companies use to store and manage their experiments, samples, instruments, and analysis. Benchling has been around for since 2012. In October 2025, it launched Benchling AI, an intelligence layer with a chat interface, backed by an agent, that helps scientists find data, design experiments, and write reports. Nick came to Benchling through its acquisition of Sphinx Bio, the analysis startup he founded. In this conversation, Nick walks through what it takes to build agents for scientific work, and where the playbook from coding agents holds up and where it breaks down.
We also discuss:
• Why Benchling invests so heavily in getting clean data upfront
• How they cross-check answers between models to get more out of each one
• Why and how Benchling leans on production traces
• Where AI actually helps science today, and where it still gets stuck
• Why understanding LLMs is closer to biology than software engineering
Timestamps:
00:00 Intro
01:22 What Benchling AI is, and the 14-year data platform underneath it
04:36 Why a decade of structured data is a core advantage
05:57 The architecture under the hood
08:28 Similarities and differences compared to a coding harness
11:14 Benchling’s multi-agent architectures
14:36 Dealing with verifiable vs non-verifiable tasks
16:19 Doing evals when clean benchmarks aren’t possible
18:13 Context engineering: SQL vs. file-based harnesses
22:11 Memory: agents that create and update their own skills
25:30 What user education for scientists looks like
30:33 Why understanding LLMs is closer to biology than software
33:28 When will agents discover a novel cure for disease?
44:58 The future of harnesses in science
48:13 Why fine-tuning on biology hasn't beaten frontier models
References:
• Agent Skills (Claude Docs): https://docs.claude.com/en/docs/agents-and-tools/agent-skills/overview
• Benchling’s Deep Research Agent: https://www.benchling.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: Agent Foundations
View skill →Related Reads
📰
📰
📰
📰
The agent fixed one hang, then immediately wrote another.
Dev.to · Bryce Darling
The AI Gold Rush’s Weirdest Side Effect: Everyone Now Builds the Same App
Medium · AI
Your agent's output is invalid and you'll find out the expensive way
Dev.to · Foxy_Grandpa
Decision fatigue in the age of infinite AI-generated options
Medium · AI
Chapters (15)
Intro
1:22
What Benchling AI is, and the 14-year data platform underneath it
4:36
Why a decade of structured data is a core advantage
5:57
The architecture under the hood
8:28
Similarities and differences compared to a coding harness
11:14
Benchling’s multi-agent architectures
14:36
Dealing with verifiable vs non-verifiable tasks
16:19
Doing evals when clean benchmarks aren’t possible
18:13
Context engineering: SQL vs. file-based harnesses
22:11
Memory: agents that create and update their own skills
25:30
What user education for scientists looks like
30:33
Why understanding LLMs is closer to biology than software
33:28
When will agents discover a novel cure for disease?
44:58
The future of harnesses in science
48:13
Why fine-tuning on biology hasn't beaten frontier models
🎓
Tutor Explanation
DeepCamp AI