Perplexity.AI Reduced Tool-Call Failures by 21 Percent
The company deployed a self-distillation training method to improve accuracy for multi-source search queries.
Updated on Sept. 22, 2026 in Artificial Intelligence

Live Poll
Do you trust that AI search tools are becoming more reliable in providing accurate information?
Perplexity.AI has implemented a new self-distillation training method that reduced tool-call failures by 21.2 percent in live testing. This development is currently deployed within the company's production environment.
Why it matters
Reducing errors in tool use lowers compute costs and latency while increasing reliability, serving as a key competitive differentiator for AI-driven search engines. This update addresses the industry-wide challenge of managing multi-source information retrieval more effectively.
The training pipeline utilizes the DART-SD framework to optimize tool calls, achieving a 21.2 percent failure reduction versus previous iterations. The model’s performance is explicitly measured against the FRAMES benchmark for multi-source question handling.
The players
Perplexity.AI
An AI search company that focuses on real-time information retrieval by integrating large language models with external tool-use capabilities.
The details
The training process begins with supervised fine-tuning, where the model learns from curated examples of accurate tool usage. Subsequently, the system undergoes on-policy reinforcement learning—a technique where the model learns by interacting with its environment—incorporating actual user corrections and real-world tool error data to refine decision-making. By automating this feedback loop, the model improves its ability to discern when and how to access external information sources.
Timeline
September 22, 2026: The improvement was detailed in a report.
The Tech Race
This integration follows the standards established by the DART-SD training framework to optimize agentic behavior. It marks a significant effort by the company to outperform standard large language models in multi-source accuracy using the FRAMES benchmark.
Users will experience fewer failed queries when the search tool attempts to aggregate information from multiple external sources. The improvement is already integrated into the current platform, requiring no action from the end user to benefit from the reduced error rate.
The takeaway
The deployment of this self-distillation method highlights a shift toward using real-world error data to tune search agents for reliability. Observers should track future updates to the FRAMES benchmark scores to determine if these gains hold across larger, more complex question sets.
Further reading
For more information on how current research shapes search performance, visit Artificial Intelligence.
Source note: This article includes information reported by Crypto Briefing.
Live Poll
Do you trust that AI search tools are becoming more reliable in providing accurate information?









