아래로 당겨서 돌아가기
You're sleeping on Devstral Small 2 - 24B Instruct

You're sleeping on Devstral Small 2 - 24B Instruct

You're sleeping on Devstral Small 2 - 24B Instruct

Peeps are always asking which local model is best. That question is loaded and totally depends on the task you're asking of it. Llama2 for example is old but still useful for summarizing YouTube transcripts into 10 bullet points. I don't code with it obviously, but it works well for that. So, I built my own benchmarking tool to test local models on my client codebases. SWE Bench and similar tools test only Python gate-based tasks. They do not care whether the model writes slop or creates addit