Perplexity.AI has deployed a new self-distillation training method that cut tool-call failures by 21.2% in live testing. For a company whose entire product depends on AI models correctly fetching, parsing, and synthesizing information from external tools, that’s the kind of improvement that separates a useful answer from a hallucinated one.
The reduction targets a specific and genuinely annoying problem in modern AI systems. When a language model needs to call an external tool, like a search engine, a calculator, or a code interpreter, it can fail in dozens of ways. It might call the wrong tool, format the request incorrectly, misinterpret the response, or call a tool when it didn’t need to at all. Every one of those failures degrades the final answer a user sees.
How self-distillation fixes broken tool calls
The approach aligns with a broader framework known as DART-SD that has been gaining traction in the AI research community. DART-SD applies localized self-distillation specifically designed for multi-turn tool-calling agents. The key insight is that it corrects failures without penalizing the valid reasoning steps the model took along the way.
Perplexity’s implementation fits into a two-stage training pipeline. First comes supervised fine-tuning, where the model learns from curated examples of correct tool use. Then comes on-policy reinforcement learning, where the model practices in something closer to real-world conditions and gets rewarded for using tools efficiently and accurately.
What makes this particularly interesting is that the reinforcement learning stage incorporates real-world user corrections and actual tool errors. The model isn’t just training on clean, idealized examples. It’s learning from the messy reality of how tools actually behave in production, including when they time out, return unexpected formats, or simply break.
Why tool-call efficiency matters more than raw intelligence
Every unnecessary tool call costs compute, adds latency, and introduces another opportunity for error. Perplexity has been pursuing a deliberate strategy of achieving higher accuracy with fewer tool calls. Their work on the FRAMES benchmark, which evaluates how well AI systems handle complex multi-source questions, has focused specifically on maintaining or improving scores while operating under controlled tool budgets.
The competitive landscape for AI search
Perplexity occupies an increasingly contested space. Google has aggressively expanded its AI Overviews feature, OpenAI has integrated web search into ChatGPT, and a growing list of startups are building search-augmented AI products. In this environment, the quality of tool use, meaning how reliably a model can search, retrieve, and synthesize information, is a core differentiator.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

57 minutes ago
32






English (US) ·