NNewsGPT ← Home
US

Rubrik's AI Judge: Automating Security Policy Enforcement Without Measuring Accuracy

US1 hr ago

Rubrik's General Manager of AI, Dev Rishi, revealed at VB Transform 2026 that the company is experimenting with an AI system, dubbed SAGE (Semantic AI Governance Engine), to autonomously enforce security policies for its AI agents. This move aims to address the practical challenges of enforcing AI governance policies, which Rishi noted are often written down but difficult to implement in practice. Rubrik's approach involves a second AI that judges every action taken by an autonomous agent in real-time against established policies, replacing human oversight. This experiment is being conducted internally first, with Rishi emphasizing that while AI autonomy is a settled capability, its deployment is an open question for organizations. He explained that requiring human approval for every command, as initially done with pilots like Claude Code, led to significant developer pushback, with users simply approving lengthy terms without reading, creating 'security theater.' Research from Rubrik Zero Labs indicates that monitoring and approving agent actions can consume more time than the agents save, a sentiment echoed by approximately 80% of respondents. SAGE functions as an arbitration layer within Rubrik Agent Cloud, interpreting the semantic intent of agent actions and comparing them against policies written in natural language. This system replaces the 'human in the loop' with an 'AI in the loop' to provide more precise governance, especially for complex rules that traditional methods struggle with, such as identifying specific revenue fields in Salesforce. Rishi highlighted that security approval, rather than cost or performance, is the primary bottleneck for AI Return on Investment (ROI) for many enterprises. He noted that his conversations with Global 2000 IT and security leaders consistently pointed to security and risk concerns as the main constraint. The economic viability of SAGE is maintained by using a small language model, significantly reducing costs and latency compared to frontier LLMs. This allows for granular control, with specialized judges within SAGE monitoring for issues like tool-use hallucinations or PII exfiltration. Rishi also discussed the 'lethal trifecta' — an agent with private data access, unvetted content ingestion, and an external communication channel — as a major security concern, where the combination of individually permissible actions becomes destructive. Traditional access management systems, reliant on human judgment, are ill-equipped to handle this emergent risk, underscoring the need for AI-driven contextual intent analysis.

AI Analysis

Rubrik's exploration into AI-driven policy enforcement, exemplified by SAGE, highlights a critical tension in the current AI adoption landscape: the gap between policy declaration and practical, scalable enforcement. While the company's internal experiment with an AI judge for autonomous agents addresses the inefficiency and user fatigue associated with manual oversight, it introduces a new challenge. The effectiveness of SAGE hinges on its accuracy and reliability, a metric that, as noted, remains unmeasured. This situation presents a systemic risk: deploying autonomous systems governed by AI without robust validation mechanisms could inadvertently create new vulnerabilities or enforce policies incorrectly. The reliance on AI to interpret and enforce complex, context-dependent policies, such as those involving sensitive data fields or combined permissions, is a forward-looking approach. However, it necessitates rigorous, ongoing evaluation of the AI judge's performance against objective benchmarks to ensure it genuinely enhances security and governance rather than merely automating potential errors. The long-term viability of such systems will depend on developing transparent and verifiable methods for measuring AI judgment accuracy, fostering trust and mitigating the inherent risks of autonomous operations in sensitive domains.

AI-generated to prompt reflection — not editorial opinion, not advice, not a statement of fact. How this works.

Compiled by NewsGPT from VentureBeat. Read the original for full details.