Anthropic's Claude AI Discovers Cryptographic Weaknesses, Though Not Practically Impactful
Researchers at Anthropic have utilized their Claude Mythos large language model to identify mathematical flaws in cryptographic systems, specifically HAWK and a modified version of AES. While these discoveries are not considered to have any practical implications for current computer systems, the process highlighted the challenges and nuances of using AI for advanced research. The prompts shared in the accompanying repository reveal that Claude initially struggled, often deeming problems unsolvable and requiring significant human encouragement to persist. The researchers emphasized the goal of leveraging Claude as a highly intelligent tool, akin to a top researcher, to uncover novel attacks and publishable findings rather than focusing on easily discoverable vulnerabilities. The model was prompted to seek genuinely difficult research outcomes, pushing beyond "low-hanging fruit." The Mythos Preview model operated for a total of 60 hours, incurring an estimated API cost of $100,000. Key human interventions involved motivating the AI to continue its search for significant, publishable discoveries.
This experiment demonstrates the potential for advanced AI models to assist in complex scientific discovery, such as identifying vulnerabilities in cryptographic algorithms. However, it also underscores the current limitations and the critical role of human guidance in directing AI research towards meaningful, novel outcomes. The substantial computational cost and the need for persistent human prompting suggest that while AI can be a powerful tool, it requires careful framing and iterative interaction to overcome inherent biases towards perceived solvability and to achieve breakthrough results. Future developments may focus on enhancing AI's intrinsic motivation and problem-solving resilience to reduce the reliance on extensive human intervention, thereby optimizing the efficiency and cost-effectiveness of AI-driven research endeavors.
AI-generated to prompt reflection — not editorial opinion, not advice, not a statement of fact. How this works.