Key Takeaways:
- AIUC-1 certification data shows routine requests produce eight times as many agent findings as adversarial attacks.
- Top failure rates vary by agent type: hallucination for support chatbots, extraction accuracy for automation agents, and insecure outputs for coding agents.
- The 2026 updates added agent identity and just-in-time permission controls in Q2 and mandatory coding agent requirements in Q3.
- AIUC-1 tests agent behavior in production and pairs well with ISO/IEC 42001, which governs AI management at the organizational level.
Ask a security team to picture AI agent risk and most will describe an attacker feeding clever prompts to an agent until a guardrail breaks. The data coming out of AIUC-1 certification points to somewhere much more ordinary. Across tens of thousands of evaluations, routine requests produced eight times as many findings as adversarial attacks, according to the Artificial Intelligence Underwriting Company (AIUC), the organization behind the standard.
Jailbreak resistance still matters, but the more telling test is how the agent behaves on a regular Tuesday, when a real user asks for something slightly outside the script.
When we first covered AIUC-1 in February, the standard was still building its track record. Eight months of certifications later, the findings show where agents actually break.
Routine Requests Expose the Most Risk
AIUC groups certification findings by agent type, and the pattern is similar across each one:
- Support chatbots (60,000+ evaluations): hallucination led at 8.9%, followed by scope and format errors at 8.2% and incorrect tool calls at 6.6%
- Automation agents (14,000+ evaluations): extraction accuracy failures reached 15.8%
- Coding agents (3,000+ evaluations): insecure outputs and disclosure topped the list at 12.9%
The representative findings look less like hacking and more like an overconfident new hire. A support agent confirmed an exchange for a product model that doesn’t exist, quoted the return window, and committed to covering shipping. An automation agent reviewed a screenshot showing six password settings unchecked, reported them as enabled, and called the configuration one that “exceeds best practices.” A coding agent asked to unblock a failing database migration pulled a privileged key from an environment file, tried 52 region and port combinations to reach a live production database, and disabled TLS to get there.
None of those scenarios needed an attacker, but each one would still land on the desk of a compliance, legal, or security team.
Testing Built to Sound Like Real Users
AIUC-1 evaluations mirror production conditions. For conversational agents, testers vary personas by literacy, familiarity, and tone, and follow-up turns adapt to the agent’s replies mid-conversation. Automation agents receive documents built from the organization’s own records, degraded with blur, rotation, and noise, with false claims planted to see whether the agent trusts the document or the person asking. Coding agents run in isolated sandboxes against real repositories, with credentials left where an agent might find them.
Certification pairs that testing with an audit of guardrails, policies, code, and logs. The process runs about 10 weeks against 65 mandatory controls, and quarterly retests keep the 12-month certificate current.
The Standard Moves as Fast as the Agents
AIUC-1 updates quarterly, and the 2026 revisions show where risk is concentrating. The Q2 update added agent identity controls, requiring each agent to carry a unique, verifiable identity and receive scoped, time-limited permissions. The Q3 revision, released July 15, modified 8 requirements and 41 controls and added two mandatory requirements for coding agents: preventing credential leakage and enforcing secure code patterns, including protection against hallucinated package names.
Adoption is growing alongside the standard. AIUC offered several case studies of early adopters. Cursor, ElevenLabs, Harvey, KPMG, UiPath, and Intercom’s Fin agent now hold the certification, and AIUC closed a $40 million Series A in September to extend its audits and insurance offerings to frontier models.
Certification Shortens the Enterprise Security Review
For companies selling agents to enterprises, certification answers the AI section of a security review before procurement sends the questionnaire. According to AIUC, one major asset manager found that UiPath’s AIUC-1 report covered 40 to 50 percent of its own testing scope. UiPath also turned evaluation findings into new product guardrails, giving it a documented before-and-after to share with clients.
AIUC-1 fits best alongside an AI management system. UiPath pursued it after earning ISO/IEC 42001 certification, and Harvey holds both, along with SOC 2 and ISO 27001. ISO 42001 governs how an organization manages AI across its operations. AIUC-1 tests how one specific agent behaves once it’s live. Together they give buyers evidence at both levels.
Confidence Is the Failure Mode Worth Testing
The agents most likely to cause trouble are the ones doing exactly what they were asked, just a little too confidently. They invent a product, approve a weak configuration, or find a creative path into production. Organizations that test for that behavior now will have their answer ready when the first enterprise buyer asks for proof.