Kimi K3's costly refusal leaks its prompt, its jailbreak, and a Claude tell
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Quoting Kimi K3 simonwillison.net
Is there something I can actually help you with today? — Kimi K3 , after refusing to leak its system prompt Tags: kimi , ai-personality , generative-ai , ai , llms
References
Simon Willison — ‘Kimi K3’ (Jul 16, 2026) simonwillison.net
The request consumed 16,658 output tokens — 13,241 of which were internal reasoning tokens — costing approximately 25 cents for a single prompt.
smol.ai AI News digest / HN thread 48935342 news.smol.ai
A simple ‘hi’ resulted in an 86-token input count on the official API, suggesting the presence of an approximately 85-token hidden system prompt… the model explicitly referred to itself as ‘Claude’ or quoted Anthropic’s specific refusal verbiage.
nxcode.io — Kimi K3 Benchmarks & Coding Agent Evaluation Guide nxcode.io
Users have documented successful bypasses using a specific ‘ENI for Kimi K3’ prompt, which reportedly renders the model an ‘open book’ when applied via the system prompt area of the API.
Washington Post — coverage of open-weight safety debate vertexaisearch.cloud.google.com
Moonshot AI has acknowledged reports of ‘model jailbreaking’ but recently classified such issues as ‘out-of-scope’ for urgent remediation, viewing them as behavioral quirks rather than critical infrastructure vulnerabilities.
Grokipedia — ‘Pelican on a bicycle AI benchmark’ grokipedia.com
By mid-2026, critics noted that most frontier models could produce a ‘passable pelican,’ rendering the original prompt less effective at distinguishing between top-tier reasoning and mere memorization.