Can agents preserve performance under clinical-scale workloads?
Does lightweight orchestration outperform a single agent when simultaneous clinical tasks become numerous and heterogeneous?
Single-agent and orchestrated multi-agent designs completed retrieval, extraction, and dosing tasks in batches ranging from 5 to 80.
Across models, multi-agent accuracy remained 65.3% at 80 tasks versus 16.6% for a single agent, while using up to 65-fold fewer tokens.
At clinical scale, coordinating specialized agents with bounded responsibilities may matter more than building ever-larger prompts.