Qwen 3.8 27B

Qwen 3.8 27B

We’re on a journey to advance and democratize artificial intelligence through open source and open science. This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, Token Speed, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model. For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Qwen Cloud . In particular, Qwen3.8-27B will be available as a hosted version with more production features, e.g., 1M context length by default, official built-in tools. For more information, please refer to the Qwen3.8-27B Overview . The service is coming soon. Stay tuned for updates.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, we are pleased to introduce Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability. For streamlined integration, we recommend using Qwen3.8 via APIs. Inference efficiency and throughput vary significantly across frameworks. We recommend using the latest framework versions to ensure optimal performance and compatibility. For production workloads or high-throughput scenarios, dedicated serving engines such as SGLang, vLLM, or Token Speed are recommended. Qwen3.8 can be deployed with popular inference frameworks, e.g.: Qwen3.8 models operate in thinking mode by default, generating thinking content signified by \n... \n\n before producing the final response.

To disable thinking content and obtain a direct response, refer to the examples here . We recommend using the following sets of sampling parameters for generation: Please note that the support for sampling parameters varies according to inference frameworks. Qwen3.8 comes with official support for reasoning_effort , which can be used to adjust reasoning depth and control cost: In addition, preserve_thinking is enabled by default for all workloads for the best out-of-the-box experience. To disable preserved thinking, refer to the examples here . In multi-turn agentic tasks, lower reasoning effort does not always reduce overall task completion time. Although it may produce faster per-turn responses, it can also lead to insufficient analysis, more failures, and repeated retries, which may increase total latency and token consumption. The Chat Completions API can be used with most inference frameworks, as well as Qwen Cloud .

Sources cited: 📰 Hacker News ↗

⚡ Effects Interpreter

🌍World Economy

  • Economies far from the headline can still catch the aftershocks.
  • Markets around the world may take their cue from how this story unfolds.

🏙️Local Economy

  • Your weekly shop could get a touch dearer, or cheaper, further down the road.
  • Jobs and trade close to home may feel a soft knock-on effect.

🏦Rates & Banks

  • Any move in rates would probably come later, not overnight.
  • Interest rates and mortgage bills are unlikely to jump straight away from this alone.

❤️Health

  • Neighbours and families may feel more anxious until the dust settles.
  • Looking after mental health is worth it when headlines feel heavy.

💷Wealth

  • Nest eggs can wobble briefly before finding their footing again.
  • Long-term savers usually ride out these small bumps just fine.

🏠Housing

  • Buyers and renters may notice only a gentle drift, if anything at all.
  • Home costs usually respond later, once the bigger picture settles.
Share: 𝕏 Twitter Facebook LinkedIn WhatsApp

Editorial note: This analysis was produced by the News Effects Interpreter, an AI editorial tool that cross-references 1 independent news sources and contextualises events in terms of their real-world impact on ordinary people. Original reporting is linked above. News Effects does not alter the facts of source reports.