ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

Qwen3.8-Flash-Next 模型解析与实战部署指南:混合注意力架构、N-gram 嵌入与长上下文推理

Qwen3.8-Flash-Next 模型解析与实战部署指南:混合注意力架构、N-gram 嵌入与长上下文推理 人工智能基础模型大模型多模态【免费下载链接】Qwen3.8-Flash-Next项目地址https://ai.gitcode.com/hf_mirrors/Qwen/Qwen3.8-Flash-Next点击查看免费下载本篇技术指南以仓库根目录 README.md 为主体围绕开源权重模型Qwen3.8-Flash-Next展开系统讲解其面向 Qwen4 的混合注意力架构Gated DeltaNet Qwen Sparse Attention、Gated Residual、N-gram Embedding 与定制训练配方等核心创新并结合 config.json、generation_config.json、chat_template.jinja 等仓库配置完整给出从服务部署、OpenAI 兼容 API 调用文本/图像/视频、思考模式控制到超长上下文扩展的端到端实战方案。读完本文你将掌握该模型 125B 参数的内部结构、各配置文件的关键字段含义以及在生产环境中配置 1M 上下文与多模态推理的具体做法。一、仓库与模型定位一个实验性架构预览Qwen3.8-Flash-Next是 Qwen 系列中一个以架构创新为核心的开源权重版本。仓库本身是一个标准 Hugging Face Transformers 格式的模型仓库包含模型权重131 个分片model-00001-of-00131.safetensors至model-00131-of-00131.safetensors总大小约 360 GB见 model.safetensors.index.json 中metadata.total_size 359999963128配置与预处理文件config.json、generation_config.json、tokenizer_config.json、preprocessor_config.json、video_preprocessor_config.json对话模板chat_template.jinja 与 tokenizer 内嵌的聊天模板二者共同决定推理时的提示词拼装方式。根据 README.md这批产物与 Hugging Face Transformers、vLLM、SGLang、TokenSpeed 等主流推理框架兼容而Qwen3.8-Flash则是基于该开源版本的官方 API 版本默认提供 1M 上下文与官方内置工具二者是开源底座与生产化服务的关系。值得强调的是README 明确指出这是一个实验性预览experimental preview其架构将是后续 Qwen4 的基础——理解它的动机是不仅追求规模更追求效率在参数规模与上下文窗口不断增长的背景下架构层面的创新才是可持续进步的方向。二、四大核心架构创新HighlightsREADME 将 Qwen3.8-Flash-Next 的首发亮点归纳为四点这四点构成了理解整个 config.json 的钥匙。2.1 混合注意力Gated DeltaNet Qwen Sparse AttentionQSAGated DeltaNet线性注意力承担绝大部分层36 层的序列混合职责以 O(1) 的线性复杂度维护状态适合超长上下文下的低成本处理。Qwen Sparse AttentionQSA与早期的 Gated DeltaNet Gated Attention 配对不同新版将稀疏注意力端升级为 QSA——不再逐 token 挑选而是以微块micro-block级别进行选择性处理。这能显著降低长上下文延迟尤其适合以 Agent 为主导的现实负载。在 config.json 中该结构直接体现在layer_types字段上48 层中依次交替排列linear_attention与full_attention即3 层线性注意力 1 层全注意力的周期重复共 12 组。这与 README 中Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))的描述完全对应。2.2 Gated Residual带门控的残差流深层 LLM 训练之所以可控依赖的是带归一化的残差流。Gated Residual 在此基础上通过**逐元素、数据依赖的读门read gate与逐分支的标量写门write gate**来调制加宽后的残差流信息。其收益是层间表达更细粒度、训练稳定性保持、推理开销增量很小。在 config.json 中对应字段为hc_count: 4分支数 4与hc_lowrank: 320瓶颈秩 320并在 model.safetensors.index.json 中可看到hyper_connection_mixer、attn_hyper_connection等权重分组含block_inject_weight、input_mix_weight_down/up、hc_norm即为该门控机制的实际参数载体。2.3 N-gram Embedding参数扩展的新轴词嵌入提供了一条低计算成本、易于 offload的参数扩展路径相比 MoE 更适合显存受限的加速器。实现方式是以短 n-gram 作为索引来扩展参数规模而不牺牲质量。仓库配置中ngram_vocab_size_base: 200000002000 万 n-gram 词表README 标注为 layer 2 处的 bigrams/trigramsngram_size: 3、make_ngram_vocab_size_divisible_by: 128、split_ngram_parts: 128ple_layer_ids: [2]仅在 layer 2 引入、ple_embed_dim: 2560、ple_conv_kernel_size: 4。README 的模型总参数构成即为125B 总参数 6B 激活参数 51B n-gram 嵌入 4B MTP可见 n-gram 嵌入占据了大量非激活参数。2.4 定制训练配方Muon AdamW 双优化器README 说明Muon 与 AdamW 优化器被应用于特定的权重类别以最大化效率在重拟合缩放定律指导下取消了传统的 batch size warmup直接以目标 batch size 启动训练从而显著减少优化器步数并支持更大的学习率以稳健收敛。这是预训练 后训练两阶段产物得以成形的重要工程前提。三、Model Overview从 README 到 config.json 的逐项对照README 给出的模型规格非常详细下表将其与 config.json 的实际字段一一对应便于你在部署时核查配置维度README 描述config.json 对应字段/值类型Causal LM with Vision Encoderarchitectures: Qwen4ExpForConditionalGenerationmodel_type: qwen4_exp训练阶段Pre-training Post-training—激活/总参数125B 总、6B 激活、51B n-gram、4B MTP见权重索引与ngram_vocab_size_base等隐藏维度2560hidden_size: 2560Token 嵌入248320 (Padded)vocab_size: 248320层数48num_hidden_layers: 48Hidden Layout12 × (3×(DeltaNet→MoE) → 1×(QSA→MoE))layer_types: 36 个linear_attention 12 个full_attentionDeltaNet 头V 48 / QK 16头维 128linear_num_value_heads: 48、linear_num_key_heads: 16、linear_key_head_dim: 128、linear_value_head_dim: 128、linear_conv_kernel_dim: 4QSA 头Q 24 / KV 2头维 256RoPE 维 64num_attention_heads: 24、num_key_value_heads: 2、head_dim: 256、partial_rotary_factor: 0.25IndexerMQA4 查询头 1 共享键头头维 128预算 512 块 / 2048 tokenindexer_n_heads: 4、indexer_kv_heads: 1、indexer_head_dim: 128、indexer_budget: 2048、indexer_compress_ratio: 4MoE512 专家10 路由 1 共享中间维 640num_experts: 512、num_experts_per_tok: 10、moe_intermediate_size: 640、shared_expert_intermediate_size: 640Gated Residual4 分支瓶颈秩 320hc_count: 4、hc_lowrank: 320LM 输出248320 (Padded)lm_head.weight存在于权重索引中MTP1 层多步训练mtp_num_hidden_layers: 1、mtp.hybrid: true、mtp.layer_types: [full_attention]上下文长度原生 262,144可扩展至 1,000,000max_position_embeddings: 262144、rope_parameters.rope_type: default几个值得注意的细节位置编码config.json 中rope_parameters默认rope_type: default、rope_theta: 10000000、mrope_interleaved: true、mrope_section: [11, 11, 10]——这是多模态 MRoPE 配置视觉 token 与文本 token 使用不同的旋转维度段。生成默认参数generation_config.json 中do_sample: true、temperature: 1.0、top_k: 20、top_p: 0.95与 README 推荐的思考模式采样参数一致eos_token_id为[248046, 248044]|im_end|与|endoftext|。对话与思维标签tokenizer_config.json 中think//thinkid 248068/248069与工具调用标签tool_call/tool_response均为普通 token说明思维链与工具调用是模型语言能力的一部分|image_pad|248056与|video_pad|248057则对应视觉占位符与 config.json 的image_token_id/video_token_id一致。四、性能基准解读Benchmark ResultsREADME 用两张对照表给出评测数据数据为仓库声明非本文实测。阅读时请注意其评测口径语言能力对照 Qwen3.8-27B、Qwen3.7-Plus、DeepSeek-V4-Flash-0731、Claude-Opus-4.6 (Max)在 125B 总参数 / 6B 激活的前提下多项 Agentic Coding 指标取得该行最优如 DeepSWE 1.158.7、SWE-bench Pro62.5、SWE-bench Multilingual81.0Agent 类任务表现突出CoWorkBench 73.9、JobBench 55.7、Toolathlon VerifiedPass173.5Agents Last Exam 的 Pass1 为 24.3、Score 51.2通用能力方面IFBench 81.3、GPQA Diamond 91.7、LiveCodeBench v6 91.9 均为行内最优HLE 35.9、NL2Repo-Bench 48.1 略低于行内最高。视觉语言能力对照 Qwen3.8-27B、Qwen3.7-Plus、Claude-Opus-4.6 (Max)多模态 Agent 场景ClawEval-MM、RecreationBench、AndroidWorld、OSWorld 2.0、Vision2Web与通用多模态ERQA、LVBench、RealWorldQA、MathVision、CharXiv多项取得最优或并列最优。必须注意的评测条件README 脚注原文信息DeepSWE 1.1 与 SWE-bench Multilingual 使用 Claude Code / mini-SWE-agent harnesstemp1.0、top_p0.95、256K 上下文SWE-bench Pro 除 Claude-Opus-4.6 使用官方公布分数外其余模型均用 Claude Code harness 复评NL2Repo-Bench 为避免 reward hacking禁用了访问特定仓库的 Bash 命令pip download、pip install、git cloneCoWorkBench 与 RecreationBench 为内部自研基准HLE 由 GPT-4o 判分行内最优以加粗显示。这些细节提示我们不同模型、不同 harness、不同采样设置下的分数不可直接横比引用时应保留脚注条件。五、快速上手部署与 API 调用Quickstart5.1 服务端部署ServingREADME 明确给出建议推理效率与吞吐在不同框架间差异显著生产负载或高吞吐场景强烈推荐使用专用服务引擎SGLang、KTransformers、vLLM并优先使用各框架最新版本以保证性能与兼容性。相关部署手册Cookbook / Recipe由各框架官方提供。该模型也支持 Hugging Face Transformers 直接加载推理。5.2 API 调用的通用准备Chat Completions API 可被绝大多数推理框架使用。以 OpenAI 兼容客户端为例先安装并配置环境变量pip install -U openai # Set the following accordingly export OPENAI_BASE_URLhttp://localhost:8000/v1 export OPENAI_API_KEYEMPTY模型名使用Qwen/Qwen3.8-Flash-Next。5.3 思考模式与采样参数重要Qwen3.8-Flash-Next默认以思考模式运行会在最终回复前生成以think\n.../think\n\n标记的思考内容。README 推荐两套采样参数模式temperaturetop_ptop_kmin_ppresence_penaltyrepetition_penalty思考模式Thinking1.00.95200.00.01.0直答模式Instruct / Non-Thinking0.70.80200.01.51.0注意各推理框架对采样参数的支持范围不尽相同。同时README 提醒在多轮 Agent 任务中降低推理强度reasoning effort不一定缩短总耗时——单轮响应虽快但分析不足可能导致失败重试反而增加总延迟与 token 消耗。模型通过三个参数控制思考行为enable_thinking、preserve_thinking、reasoning_effort。这三个参数也正是 chat_template.jinja 中模板渲染所读取的核心开关enable_thinking默认 true控制是否生成思考块当显式置为 false 时模板在生成提示词末尾仍会注入空思考块think\n\n/think\n\n以保证推理引擎行为一致reasoning_effort默认xhigh可选xhigh/medium/low模板据此注入对应的推理指令文本xhigh 要求仔细思考、校验关键假设、权衡替代方案low 要求思考简短聚焦、直接给出结论preserve_thinking默认 true决定历史消息中的思考块是否保留详见下文 5.7。5.4 文本输入 流式输出含 Usagefrom openai import OpenAI # Configured by environment variables client OpenAI() messages [ {role: user, content: Write a Python function to merge two sorted linked lists.}, ] completion client.chat.completions.create( modelQwen/Qwen3.8-Flash-Next, messagesmessages, extra_body{ chat_template_kwargs: { enable_thinking: True, # on by default preserve_thinking: True, # on by default }, }, reasoning_effortxhigh, # xhigh by default; supported levels are xhigh, medium, and low streamTrue, stream_options{include_usage: True}, ) reasoning_content answer_content is_answering False print(\n * 20 Reasoning * 20 \n) for chunk in completion: if not chunk.choices: print(\nUsage:) print(chunk.usage) continue delta chunk.choices[0].delta if hasattr(delta, reasoning_content) and delta.reasoning_content is not None: if not is_answering: print(delta.reasoning_content, end, flushTrue) reasoning_content delta.reasoning_content elif hasattr(delta, reasoning) and delta.reasoning is not None: if not is_answering: print(delta.reasoning, end, flushTrue) reasoning_content delta.reasoning if hasattr(delta, content) and delta.content: if not is_answering: print(\n * 20 Answer * 20 \n) is_answering True print(delta.content, end, flushTrue) answer_content delta.content messages.append({ role: assistant, content: answer_content, reasoning_content: reasoning_content, reasoning: reasoning_content, })这段代码的要点流式 chunk 中思考内容可能通过delta.reasoning_content或delta.reasoning两个字段之一携带兼容不同框架最终把思考与答案拼回 assistant 消息回传以维持多轮对话的上下文连续性。5.5 图像输入Image Inputfrom openai import OpenAI # Configured by environment variables client OpenAI() messages [ { role: user, content: [ { type: image_url, image_url: { url: your-image-url } }, { type: text, text: The centres of the four illustrated circles are in the corners of the square. The two big circles touch each other and also the two little circles. With which factor do you have to multiply the radii of the little circles to obtain the radius of the big circles?\nChoices:\n(A) $\\frac{2}{9}$\n(B) $\\sqrt{5}$\n(C) $0.8 \\cdot \\pi$\n(D) 2.5\n(E) $1\\sqrt{2}$ } ] } ] chat_response client.chat.completions.create( modelQwen/Qwen3.8-Flash-Next, messagesmessages, ) print(Chat response:, chat_response)图像在提示词中会经由 chat_template.jinja 的render_content宏被替换为|vision_start||image_pad||vision_end|视觉占位符序列视觉特征则由 Qwen3VL 处理器完成提取。5.6 视频输入Video Inputfrom openai import OpenAI # Configured by environment variables client OpenAI() messages [ { role: user, content: [ { type: video_url, video_url: { url: your-video-url } }, { type: text, text: How many porcelain jars were discovered in the niches located in the primary chamber of the tomb? } ] } ] chat_response client.chat.completions.create( modelQwen/Qwen3.8-Flash-Next, messagesmessages, ) # When vLLM is launched with --media-io-kwargs {video: {num_frames: -1}}, # video frame sampling can be configured via extra_body (e.g., by setting fps). # This feature is currently supported only in vLLM. # # By default, fps2 and do_sample_framesTrue. # With do_sample_framesTrue, you can customize the fps value to set your desired video sampling rate. # chat_response client.chat.completions.create( # modelQwen/Qwen3.8-Flash-Next, # messagesmessages, # extra_body{ # mm_processor_kwargs: {fps: 2, do_sample_frames: True}, # }, # ) print(Chat response:, chat_response)视频默认以fps2抽帧、do_sample_framesTrue如需自定义采样率可在 vLLM 以--media-io-kwargs {video: {num_frames: -1}}启动后通过mm_processor_kwargs里的fps覆盖。视频处理对应的预处理配置见 video_preprocessor_config.jsonprocessor_class: Qwen3VLProcessor、video_processor_type: Qwen3VLVideoProcessor。5.7 直答模式与思考保留控制关闭思考Instruct 模式——模型默认先思考后回答可通过参数直接得到不思考的回答from openai import OpenAI # Configured by environment variables client OpenAI() messages [ { role: user, content: [ { type: image_url, image_url: { url: your-image-url } }, { type: text, text: Where is this? } ] } ] chat_response client.chat.completions.create( modelQwen/Qwen3.8-Flash-Next, messagesmessages, temperature0.7, top_p0.8, presence_penalty1.5, extra_body{ top_k: 20, chat_template_kwargs: {enable_thinking: False}, }, ) print(Chat response:, chat_response)若使用 Qwen Cloud 的 API除更换model外应直接传enable_thinking: False而非包在chat_template_kwargs中。关闭思考保留Disable Preserved Thinking——默认情况下模型会保留所有历史消息中的思考块形成完整的推理轨迹这保证了上下文连续性尤其适合需要决策一致性与减少冗余推理的 Agent 场景同时改善了 KV cache 利用率。若只想保留最近一条用户消息的思考块可将preserve_thinking置为 Falsefrom openai import OpenAI # Configured by environment variables client OpenAI() messages [...] chat_response client.chat.completions.create( modelQwen/Qwen3.8-Flash-Next, messagesmessages, extra_body{ chat_template_kwargs: {preserve_thinking: False}, }, ) print(Chat response:, chat_response)若使用 Qwen Cloud 的 API应直接传preserve_thinking: False而非包在chat_template_kwargs中。该开关的实现逻辑可在 chat_template.jinja 中找到模板遍历历史消息时只有满足preserve_thinking未显式关闭、或该消息位于最后一次用户查询之后时才会在 assistant 消息中恢复think.../think块。六、最佳实践Best Practices6.1 采样参数与防重复复用上文的思考模式 / 直答模式两套参数即可。对支持的框架可将presence_penalty在02之间调整以抑制无休止重复但注意较高的取值偶尔会引发语言混杂与轻微性能下降。6.2 充足的输出长度Agent 任务关键为了在 Agent 任务中发挥最佳性能建议在 1M 上下文内为内部推理与最终输出分别配置独立的 token 上限Reasoning Content思考内容最大输出 262,144 tokensFinal Response最终回答最大输出 131,072 tokens。这样既为复杂推理留足空间又保证最终交付物有充分篇幅。6.3 超长文本处理YaRN RoPE 缩放模型原生支持 262,144 tokens当输入输出总长度超过该上限时README 推荐采用YaRN这类 RoPE 缩放技术vLLM、SGLang、TokenSpeed 均已支持。开启方式有两种方式一修改模型配置文件。在 config.json 的text_config.rope_parameters中改为{ mrope_interleaved: true, mrope_section: [ 11, 11, 10 ], rope_type: yarn, rope_theta: 10000000, partial_rotary_factor: 0.25, factor: 4.0, original_max_position_embeddings: 262144 }方式二启动参数覆盖无需改文件。vLLMVLLM_ALLOW_LONG_MAX_MODEL_LEN1 vllm serve ... --hf-overrides {text_config: {rope_parameters: {mrope_interleaved: true, mrope_section: [11, 11, 10], rope_type: yarn, rope_theta: 10000000, partial_rotary_factor: 0.25, factor: 4.0, original_max_position_embeddings: 262144}}} --max-model-len 1000000SGLangSGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN1 python -m sglang.launch_server ... --json-model-override-args {text_config: {rope_parameters: {mrope_interleaved: true, mrope_section: [11, 11, 10], rope_type: yarn, rope_theta: 10000000, partial_rotary_factor: 0.25, factor: 4.0, original_max_position_embeddings: 262144}}} --context-length 1000000TokenSpeedTOKENSPEED_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN1 tokenspeed serve ... --hf-overrides {text_config: {rope_parameters: {mrope_interleaved: true, mrope_section: [11, 11, 10], rope_type: yarn, rope_theta: 10000000, partial_rotary_factor: 0.25, factor: 4.0, original_max_position_embeddings: 262144}}} --max-model-len 1000000重要注意事项README 原文要点主流开源框架实现的是静态 YaRN缩放因子不随输入长度变化可能对较短文本的性能有影响因此仅在确实需要处理长上下文时才修改rope_parameters建议按实际场景调整factor例如若应用典型上下文为 524,288 tokens则把factor设为 2.0 更合适。6.4 长视频理解优化为了兼顾纯文本与图像的推理效率发布版 video_preprocessor_config.json 中的size.longest_edge被保守地配置为 25165824。README 建议在需要小时级长视频的高帧率采样时将longest_edge调至469,762,048对应约 224k 视频 token以获得更优性能例如{longest_edge: 469762048, shortest_edge: 4096}也可通过引擎启动参数覆盖默认值实现细节参见 vLLM / SGLang 相关改动。6.5 引用Citation若你的工作参考了本模型README 提供了两条 BibTeX 记录一条对应架构设计技术报告《On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability》另一条对应本模型的技术博客《Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency》可按需引用。七、总结Qwen3.8-Flash-Next 是理解下一代高效 LLM 架构的绝佳样例Gated DeltaNet 线性注意力与 QSA 稀疏注意力的混合布局config.json 中 36 层线性 12 层稀疏交替、Gated Residual 门控残差hc_count: 4、2000 万词表的 N-gram 嵌入ngram_vocab_size_base: 20000000以及Muon/AdamW 定制训练配方共同支撑起 125B 总参、6B 激活的高性价比形态。在工程侧通过enable_thinking/preserve_thinking/reasoning_effort三个参数其语义与 chat_template.jinja 的模板逻辑一一对应即可精细控制思考行为配合 YaRN 配置可将原生 262,144 的上下文扩展到 1M再结合图像、视频的多模态输入能力完全具备承载长程 Agent 任务的生产条件。建议读者在部署时结合本仓库的配置文件逐项核对并优先采用最新版本的专用推理框架以获得最佳吞吐与兼容性。赞分享人工智能基础模型大模型多模态【免费下载链接】Qwen3.8-Flash-Next项目地址https://ai.gitcode.com/hf_mirrors/Qwen/Qwen3.8-Flash-Next点击查看免费下载相关推荐Simple Live 跨平台直播聚合全解析B站、斗鱼、虎牙、抖音一个应用看完Simple Live 跨平台直播聚合全解析B站、斗鱼、虎牙、抖音一个应用看完 Simple Live 是一个用 Flutter 写的直播聚合应用把哔哩哔哩音视频直播Qwen3.8-Flash-Next混合注意力深度解析Gated DeltaNetQSA组合为何能大幅降低长上下文延迟Qwen3.8 Flash Next混合注意力深度解析Gated DeltaNetQSA组合为何能大幅降低长上下文延迟 Qwen3.8 Flash Nex人工智能基础模型大模型多模态SpringBoot应用监控与管理Actuator与SpringBoot Admin在tech-pdai-spring-demos中的应用SpringBoot应用监控与管理Actuator与SpringBoot Admin在tech pdai spring demos中的应用 SpringBoo创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
返回列表