Pull down to go back
156 Landing Page Tests: Why Detailed Design Prompts Actually Made Gemma 4 31B Worse (Not Better)

156 Landing Page Tests: Why Detailed Design Prompts Actually Made Gemma 4 31B Worse (Not Better)

156 次登陸頁面測試:詳細設計提示為什麼讓 Gemma 4 31B 表現更差(不是更好)

A developer tested Gemma 4 31B (a smaller AI model) to generate landing pages for a luxury real-estate CRM using 52 different system prompts. The surprising finding: prompts packed with detailed "design heuristics" and rules performed *worse* than giving the model almost no instructions at all. They used temperature 0.7, generated 3 samples per persona, and kept everything in a single HTML file with inline styles. The reason they picked a smaller model instead of GPT-4 or Claude Opus? Bigger models are already so good that testing different prompts becomes meaningless—you need a model with actual room to improve to see what actually matters. This is the kind of counterintuitive result that makes you rethink how you write prompts.