AI在线 AI在线

Patronus AI Launches Percival: One-Minute Diagnosis of Hidden Faults in Hundred-Step Agent Chains

As enterprises increasingly deploy autonomous AI agent systems, the demand for monitoring and debugging these complex systems is rapidly growing. Today, AI security company Patronus AI, headquartered in San Francisco, released its latest product, Percival, a monitoring platform capable of automatically identifying fault patterns in AI agent systems and providing repair recommendations."Percival is the industry's first intelligent agent that can automatically track agent trajectories, identify complex faults, and systematically output repair suggestions," said Anand Kannappan, CEO and co-founder of Patronus AI, in an exclusive interview with VentureBeat.Solving the Real-World Challenges of "Uncontrollable" AI AgentsDifferent from traditional machine learning, AI agents can autonomously execute large-scale operation processes involving multiple stages.

As enterprises increasingly deploy autonomous AI agent systems, the demand for monitoring and debugging these complex systems is rapidly growing. Today, AI security company Patronus AI, headquartered in San Francisco, released its latest product, Percival, a monitoring platform capable of automatically identifying fault patterns in AI agent systems and providing repair recommendations.

"Percival is the industry's first intelligent agent that can automatically track agent trajectories, identify complex faults, and systematically output repair suggestions," said Anand Kannappan, CEO and co-founder of Patronus AI, in an exclusive interview with VentureBeat.

Solving the Real-World Challenges of "Uncontrollable" AI Agents

Different from traditional machine learning, AI agents can autonomously execute large-scale operation processes involving multiple stages. However, it is precisely this "multi-step autonomy" that makes fault debugging extremely challenging: a small error in the early stage may evolve into a serious deviation in subsequent processes, and multi-agent collaborative scenarios further exacerbate this complexity.

Percival is designed to address this pain point, capable of identifying over 20 common faults across four major categories, including reasoning errors, execution errors, planning misalignments, and domain-specific errors. More importantly, it is not a "post-hoc" solution but actively monitors the entire agent trajectory, possessing "contextual memory" to understand the ins and outs of errors in specific contexts.

"Percival itself is also an AI agent, so unlike traditional evaluators, it does not make static judgments but can track and learn fault evolution paths at the system level," said Darshan Deshpande, a researcher at Patronus.

Holographic Projection Robot Design (2)

Image source note: Image generated by AI, licensed by Midjourney

From One Hour to One Minute: Significant Improvement in Debugging Efficiency

In practical applications, Percival has significantly improved fault analysis efficiency. Patronus stated that its early customers have compressed the time required to debug complex agent processes from about one hour to 1 to 1.5 minutes, greatly alleviating the maintenance burden on engineering teams.

To standardize evaluation capabilities, Patronus also released the TRAIL Benchmark Test (Tracking Reasoning and Agent Issue Localization). The results showed that even the strongest models currently available scored only 11% on this test. This highlights the urgent need for professional AI regulatory tools.

Enterprise Deployment and Integration: High-Complexity Agent Safety Barriers

Percival has been adopted by several clients, including Emergence AI and Nova. Satya Nitta, CEO of Emergence AI, which focuses on developing systems for "agent creation agents," said that Percival provides critical assurance for achieving controllability in large-scale autonomous systems.

Nova, on the other hand, is using Percival to build an AI-driven platform to help businesses migrate SAP systems and integrate legacy code, with their agent system processes involving hundreds of steps, far exceeding human-controlled complexity.

Percival can seamlessly integrate with mainstream frameworks such as Hugging Face Smolagents, Langchain, Pydantic AI, and OpenAI Agent SDK, covering a wide range of agent development ecosystems.

Accelerating Growth in AI Security and Regulatory Tracks

With AI technology rapidly commercializing, enterprises generate billions of lines of AI code daily. Kannappan pointed out: "Systems are becoming more and more autonomous, while human supervision capabilities are far from keeping up."

相关资讯

Patronus AI 推出 Percival:一分钟诊断百步代理链中的隐藏故障

随着企业越来越多地部署自主运行的 AI 代理系统,对这些复杂系统的监控与调试需求也迅速增长。 总部位于旧金山的 AI 安全公司 Patronus AI 今日发布了其最新产品 Percival,一个能够自动识别 AI 代理系统中故障模式并提出修复建议的监控平台。 “Percival 是业界首个可以自动追踪代理轨迹、识别复杂故障,并系统化输出修复建议的智能代理。
5/15/2025 11:01:55 AM
AI在线

Disrupting Tradition! New Multi-Agent Framework OWL Gains 17K Stars, Surpassing OpenAI to Pioneer a New Era of Intelligent Collaboration

With the rapid development of large language models (LLMs), single agents have revealed many limitations when dealing with complex real-world tasks. To address this issue, a new multi-agent framework named Workforce and an accompanying training method called OWL (Optimized Workforce Learning) were jointly introduced by institutions such as Hong Kong University and camel-ai. Recently, this innovative achievement achieved an accuracy rate of 69.70% on the authoritative benchmark test GAIA, not only breaking the record for open-source systems but also surpassing commercial systems like OpenAI Deep Research..
6/17/2025 9:03:21 PM
AI在线

占比 44%,报告称 OpenAI 的 GPT-4 充斥大量版权内容

根据 Patronus AI 近日发表的最新报告,OpenAI 的 GPT-4 模型中包含大量的版权内容,其占比达到了 44%。Patronus AI 是一家专门评估大型语言模型(LLMs)的公司,本周三发布的报告中测试了四款主流 AI 模型:OpenAI 的 GPT-4、Anthropic 的 Claude 2、Meta 的 Llama 2 以及 Mistral AI 的 Mixtral,意外的是没有谷歌的 Gemini。Patronus AI 使用 CopyrightCatcher 分析 4 款 AI 模型对主
3/8/2024 9:20:43 AM
故渊
  • 1