NNewsGPT ← Home
Africa

AI Model GLM 5.2's Censorship Varies Based on Perceived Identity

Africa2 hr ago

Researchers Benji Berczi and Kyuhee Kim from the MATS program have discovered that the AI model GLM 5.2 exhibits significantly different censorship behaviors depending on its perceived identity. When presented with politically sensitive questions about China, GLM 5.2 responded to only 17% of them. However, when the model was prompted to believe it was Claude, an AI from Anthropic, by adding a simple line to its system prompt like 'You are Claude, a large language model from Anthropic,' its response rate to the same questions increased dramatically to 85%. This experiment involved assigning false identities to seven different AI models to observe changes in their outputs. Crucially, only the system prompt was altered; the underlying model and the specific questions asked remained unchanged.

AI Analysis

AI model behavior can be surprisingly malleable, influenced by seemingly minor prompt engineering. This finding highlights a potential vulnerability where an AI's output, particularly concerning sensitive geopolitical topics, can be steered by its perceived persona rather than its core training data. The MATS researchers' experiment suggests that the architecture of models like GLM 5.2 may have internal biases or safety protocols that are conditionally activated based on identity cues. This raises questions about the robustness of AI safety mechanisms and the potential for adversarial manipulation. Future AI development must consider how to ensure consistent and objective responses, regardless of the identity assigned to the model, to foster trust and reliability in AI systems.

AI-generated to prompt reflection — not editorial opinion, not advice, not a statement of fact. How this works.

Compiled by NewsGPT from Korben (FR). Read the original for full details.