NNewsGPT ← Home
US

Expedia's AI Chief: Evaluations, Not PRDs, Define AI Product Development

US2 hr ago

Xavi Amatriain, Expedia Group's Chief AI and Data Officer, declared that "evals are the new PRD" at the VB Transform 2026 conference. He explained that desired product functionalities are now encoded through evaluations, including red teaming and security checks, before coding begins. This shift is driven by the increasing reliance on AI-assisted and AI-generated code, where all conceptualization will focus on these evaluation metrics. Amatriain previously held a significant role at Google, overseeing AI platforms for Gemini and Search, and has mentored founders of prominent AI companies like Perplexity and Scale AI. A VentureBeat survey revealed that 66% of 157 surveyed enterprises allow AI deployment without human review, yet only 5% fully trust automated evaluations, with half reporting AI agents failing real customers after passing internal tests.

Amatriain also cautioned against excessive "guardrails," arguing they can be brittle and negatively bias feedback loops, leading to incorrect learning. He views them as a "necessary evil" to be minimized. Expedia's governance strategy layers principles first, followed by processes and tools, and finally automation. These layers are enforced through risk-calibrated "agent release toll gates." He advocates for specialized AI agents composed into larger systems rather than a single monolithic AI, emphasizing that system design, not just the model, is crucial for security and functionality. Expedia maintains user control over final decisions like booking flights or hotels, deeming it a non-negotiable security measure. He also warned that future threats will increasingly originate from other AI systems, making rapid detection and remediation essential.

AI Analysis

The assertion that "evals are the new PRD" highlights a critical inflection point in AI development, moving from static documentation to dynamic, performance-based validation. This paradigm shift reflects the inherent complexities of AI, where emergent behaviors and nuanced performance metrics often defy traditional specification. The challenge lies in ensuring these evaluations are robust, unbiased, and truly predictive of real-world performance, especially as enterprises increasingly automate critical decisions. The tension between minimizing guardrails for better feedback and maintaining them for risk mitigation underscores a fundamental governance dilemma. As AI systems become more sophisticated and potentially adversarial, the focus on integrated security and rapid remediation cycles, as advocated by Amatriain, will be paramount for maintaining system integrity and user trust in the evolving technological landscape.

AI-generated to prompt reflection — not editorial opinion, not advice, not a statement of fact. How this works.

Compiled by NewsGPT from VentureBeat. Read the original for full details.