Pull down to go back
2b or not 2b? Custom LLM Scheduling Competition

2b or not 2b? Custom LLM Scheduling Competition

2b 還是不 2b?客製化大型語言模型排程競賽

Hey everyone, I just launched a Kaggle competition focused on a practical problem: resource management and reducing token costs for LLM responses. The core challenge is simple but clever—given a question from the MMLU benchmark, you don't answer it. Instead, you decide whether to use a small 2B model or a larger one. The goal is to find the sweet spot between accuracy and cost efficiency. I'm planning to add more models over time to make the decision-making more nuanced. If you're interested in optimizing AI inference costs or exploring how different model sizes perform on real benchmarks, this is worth checking out: https://www.kaggle.com/competitions/llm-scheduling-competition