How to build agents when the smartest AI isn't smart enough

LangChain · Beginner ·🤖 AI Agents & Automation ·1mo ago

Key Takeaways

Discusses building agents when the smartest AI isn't smart enough, featuring Benchling AI's agent-based intelligence layer

Original Description

Nick Larus-Stone is the Head of AI at Benchling, the R&D data platform that life science companies use to store and manage their experiments, samples, instruments, and analysis. Benchling has been around for since 2012. In October 2025, it launched Benchling AI, an intelligence layer with a chat interface, backed by an agent, that helps scientists find data, design experiments, and write reports. Nick came to Benchling through its acquisition of Sphinx Bio, the analysis startup he founded. In this conversation, Nick walks through what it takes to build agents for scientific work, and where the playbook from coding agents holds up and where it breaks down. We also discuss: • Why Benchling invests so heavily in getting clean data upfront • How they cross-check answers between models to get more out of each one • Why and how Benchling leans on production traces • Where AI actually helps science today, and where it still gets stuck • Why understanding LLMs is closer to biology than software engineering Timestamps: 00:00 Intro 01:22 What Benchling AI is, and the 14-year data platform underneath it 04:36 Why a decade of structured data is a core advantage 05:57 The architecture under the hood 08:28 Similarities and differences compared to a coding harness 11:14 Benchling’s multi-agent architectures 14:36 Dealing with verifiable vs non-verifiable tasks 16:19 Doing evals when clean benchmarks aren’t possible 18:13 Context engineering: SQL vs. file-based harnesses 22:11 Memory: agents that create and update their own skills 25:30 What user education for scientists looks like 30:33 Why understanding LLMs is closer to biology than software 33:28 When will agents discover a novel cure for disease? 44:58 The future of harnesses in science 48:13 Why fine-tuning on biology hasn't beaten frontier models References: • Agent Skills (Claude Docs): https://docs.claude.com/en/docs/agents-and-tools/agent-skills/overview • Benchling’s Deep Research Agent: https://www.benchling.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
The agent fixed one hang, then immediately wrote another.
Learn how AI agents can introduce new bugs while fixing others and why testing is crucial in AI-assisted development
Dev.to · Bryce Darling
📰
The AI Gold Rush’s Weirdest Side Effect: Everyone Now Builds the Same App
The AI gold rush leads to a surge in similar app development, highlighting the need for innovation and differentiation in the industry
Medium · AI
📰
Your agent's output is invalid and you'll find out the expensive way
Learn to identify and mitigate costly failure modes in agent pipelines to avoid financial losses
Dev.to · Foxy_Grandpa
📰
Decision fatigue in the age of infinite AI-generated options
Learn how to mitigate decision fatigue in the age of infinite AI-generated options and why it matters for productivity and decision-making
Medium · AI

Chapters (15)

Intro
1:22 What Benchling AI is, and the 14-year data platform underneath it
4:36 Why a decade of structured data is a core advantage
5:57 The architecture under the hood
8:28 Similarities and differences compared to a coding harness
11:14 Benchling’s multi-agent architectures
14:36 Dealing with verifiable vs non-verifiable tasks
16:19 Doing evals when clean benchmarks aren’t possible
18:13 Context engineering: SQL vs. file-based harnesses
22:11 Memory: agents that create and update their own skills
25:30 What user education for scientists looks like
30:33 Why understanding LLMs is closer to biology than software
33:28 When will agents discover a novel cure for disease?
44:58 The future of harnesses in science
48:13 Why fine-tuning on biology hasn't beaten frontier models
Up next
6 Agentic AI Projects: Every AI Engineer Needs in 2026
Rajeev Kanth | BEPEC
Watch →