wiki / raw / simon-willison-qwen3-8-flash-next-2026
Qwen3.8-Flash-Next
loading…
Original source: https://simonwillison.net/2026/Aug/26/qwen38-flash-next/ SHA256: d349241c107e31cb5b6ad9cb84af0fa5de9c1e59210084a1d2f9055e4629200e
Qwen3.8-Flash-Next
Qwen3.8-Flash-Next Simon Willison’s Weblog Subscribe Sponsored by: Greptile — The Al code reviewer that runs your code. Catch bugs that only show up at runtime. Try it for free 26th August 2026 - Link Blog Qwen3.8-Flash-Next (via) Another open weights model from Qwen. This one is “a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4”. It’s pretty big: 125B tokens, but only 6B active which means it gets a significant performance boost. I’ve been trying it out on a DGX Spark using these Unsloth quantized models. I’m still exploring the model - so far I’ve tried the 72.5GB UD-IQ1_S one (producing these pelicans) and the 78.9GB UD-Q2_K_XL (producing these). My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL: Posted 26th August 2026 at 11:52 pm Recent articles Conceptual integrity and counting lines of code - 19th August 2026 Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026 This is a link post by Simon Willison, posted on 26th August 2026. ai 2,205 generative-ai 1,954 llms 1,921 qwen 61 pelican-riding-a-bicycle 136 ai-in-china 107 nvidia-spark 6 Monthly briefing Sponsor me for $10/month and get a curated email digest of the month’s most important LLM developments. Pay me to send you less! Sponsor & subscribe Disclosures Colophon © 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026