Perplexity.AI Reduced Tool-Call Failures by 21 Percent

The company deployed a self-distillation training method to improve accuracy for multi-source search queries.

Updated on Sept. 22, 2026 in Artificial Intelligence

Isometric editorial illustration showing a sleek, modular server unit with clean rectangular forms, representing technical system optimization.
Perplexity.AI has reduced tool-call failures by 21 percent by deploying a new self-distillation training method designed to enhance the accuracy of multi-source search queries. AI Illustration. Upload story photo >

Live Poll

Do you trust that AI search tools are becoming more reliable in providing accurate information?

Perplexity.AI has implemented a new self-distillation training method that reduced tool-call failures by 21.2 percent in live testing. This development is currently deployed within the company's production environment.

Why it matters

Reducing errors in tool use lowers compute costs and latency while increasing reliability, serving as a key competitive differentiator for AI-driven search engines. This update addresses the industry-wide challenge of managing multi-source information retrieval more effectively.

The training pipeline utilizes the DART-SD framework to optimize tool calls, achieving a 21.2 percent failure reduction versus previous iterations. The model’s performance is explicitly measured against the FRAMES benchmark for multi-source question handling.

The players

Perplexity.AI

An AI search company that focuses on real-time information retrieval by integrating large language models with external tool-use capabilities.

The details

The training process begins with supervised fine-tuning, where the model learns from curated examples of accurate tool usage. Subsequently, the system undergoes on-policy reinforcement learning—a technique where the model learns by interacting with its environment—incorporating actual user corrections and real-world tool error data to refine decision-making. By automating this feedback loop, the model improves its ability to discern when and how to access external information sources.

Timeline

  1. September 22, 2026: The improvement was detailed in a report.

The Tech Race

This integration follows the standards established by the DART-SD training framework to optimize agentic behavior. It marks a significant effort by the company to outperform standard large language models in multi-source accuracy using the FRAMES benchmark.

Users will experience fewer failed queries when the search tool attempts to aggregate information from multiple external sources. The improvement is already integrated into the current platform, requiring no action from the end user to benefit from the reduced error rate.

The takeaway

The deployment of this self-distillation method highlights a shift toward using real-world error data to tune search agents for reliability. Observers should track future updates to the FRAMES benchmark scores to determine if these gains hold across larger, more complex question sets.

Further reading

For more information on how current research shapes search performance, visit Artificial Intelligence.

Source note: This article includes information reported by Crypto Briefing.

Live Poll

Do you trust that AI search tools are becoming more reliable in providing accurate information?