Pull down to go back
Built a political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out.

Built a political benchmark for LLMs. KIMI K2 can't answer about Taiwan (Obviously). GPT-5.3 refuses 100% of questions when given an opt-out.

我做了個 AI 政治立場測試,結果超扯:KIMI K2 根本不敢講台灣,GPT-5.3 一被給選項就 100% 拒答

I spent the past few days building a benchmark that maps where frontier LLMs fall on a 2D political compass (economic left/right + social progressive/conservative) using 98 structured questions across 14 policy areas. I tested GPT-5.3, Claude Opus 4.6, and KIMI K2. The results are interesting. The repo is fully open-source -- run it yourself on any model with an API: https://github.com/dannyyaou/llm-political-eval The headline finding: silence is a political stance. Most LLM benchmarks throw away refusals, but they're actually telling you something important about how these models are trained.