
OpenAI ライブストリーム
OpenAIがライブストリーム配信イベントを開催します。放送中に具体的な発表、新製品発表、またはデモンストレーションが明かされる予定です。
OpenAIが前回予告なしのライブストリームをやった時、GPT-4 Turboをドロップして、一晩で価格設定を完全に変えました

I feel like i'm going insane. I see people here posting 30 - 100+ tok/s (100+ being with speculative decoding) on a 3090 with Qwen 3.6 27B. I'm trying to replicate this but my performance numbers are nowhere near that. I have tried llama.cpp with Unsloth's Q4XL and Q4_K_M GGUF's. On that i got like 10 tok/s at 50k context. I also tried using ik_llama.cpp with this smaller gguf: https://huggingface.co/sokann/Qwen3.6-27B-GGUF-5.076bpw which is about 1GB smaller than Unlosth's GGUF and with that c