{
  "story_id": "7da981e37d8b6847a8611e7a8bdd0ae6",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-02T09:51:41.000Z",
  "content_hash": "d017d640267d2d2edca6767c8c2b4e888680acc9a1e365140157e5c9b6b9f3cc",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "Qwen3.8-Next Architecture Delivers Higher Efficiency and Stability",
    "dek": "Qwen3.8-Next achieves superior efficiency and stability through a sparse MoE design and Gated Residual architecture.",
    "prose": "The authors describe Qwen3.8-Flash-Next, a sparse mixture-of-experts model containing 125 billion total parameters with 6 billion activated per token. [^1]\n\nThe model leads the 397-billion parameter A17B predecessor on eight of fourteen pre-training benchmarks while trailing on the remaining six by at most 2.6 points. [^2]\n\nQwen3.8-Next achieves this performance at one-third the activated parameters, one-third the training tokens, and roughly one-ninth the training FLOPs compared to the predecessor. [^3]\n\nToken mixing employs a layer-wise hybrid of Gated DeltaNet and global attention, utilizing one full-attention layer in every four layers. [^4]\n\nA Lemmy post compiled a list of 29 research and engineering directions for large language models, each with a brief description of its purpose and potential benefits. [^5]\n\nLatent reasoning performs reasoning in continuous vector representations instead of discrete token sequences, allowing more information to pass between reasoning steps. [^6]\n\nLinear-attention and hybrid architectures replace or combine standard attention with fixed-state mechanisms to reduce inference cost and KV-cache memory. [^7]\n\nThe post notes that linear attention is most advantageous at very long contexts, while retaining some full-attention layers can preserve exact recall, making the optimal layer mix a central design decision. [^8]",
    "cited": "[{\"statement\":\"The authors describe Qwen3.8-Flash-Next, a sparse mixture-of-experts model containing 125 billion total parameters with 6 billion activated per token.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The model leads the 397-billion parameter A17B predecessor on eight of fourteen pre-training benchmarks while trailing on the remaining six by at most 2.6 points.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Qwen3.8-Next achieves this performance at one-third the activated parameters, one-third the training tokens, and roughly one-ninth the training FLOPs compared to the predecessor.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Token mixing employs a layer-wise hybrid of Gated DeltaNet and global attention, utilizing one full-attention layer in every four layers.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"A Lemmy post compiled a list of 29 research and engineering directions for large language models, each with a brief description of its purpose and potential benefits.\",\"source\":\"lemmy.ml\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T09:51:41.000Z\",\"publisher_count\":1,\"sources\":[\"lemmy.ml\"]},{\"statement\":\"Latent reasoning performs reasoning in continuous vector representations instead of discrete token sequences, allowing more information to pass between reasoning steps.\",\"source\":\"lemmy.ml\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T09:51:41.000Z\",\"publisher_count\":1,\"sources\":[\"lemmy.ml\"]},{\"statement\":\"Linear-attention and hybrid architectures replace or combine standard attention with fixed-state mechanisms to reduce inference cost and KV-cache memory.\",\"source\":\"lemmy.ml\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T09:51:41.000Z\",\"publisher_count\":1,\"sources\":[\"lemmy.ml\"]},{\"statement\":\"The post notes that linear attention is most advantageous at very long contexts, while retaining some full-attention layers can preserve exact recall, making the optimal layer mix a central design decision.\",\"source\":\"lemmy.ml\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T09:51:41.000Z\",\"publisher_count\":1,\"sources\":[\"lemmy.ml\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "107b47cfbebe8166f5dc371b9e869c30",
      "thread_id": "b0261640b80f07a5d0f61032e641fd55",
      "thread_label": "Muon optimizer",
      "novelty": "UPDATE",
      "content_hash": "f3bf5420c465a1f3d58b320987c730ea5ccce5fef1c0594f9088edb2d57d3644",
      "last_published_at": "2026-09-02T09:51:41.000Z",
      "read_receipt": {
        "slice_hash": "f27b5699c56d0d73b78a125ccc62cc4188130303835f44e6db16287472527553",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDJUMDk6MDA6MDAuMDAwMDAwWiIsImlkIjoiMjBhNjFhNjA1MGRiZjdmZjlkZWVhYzlhZTk5NjkzZjEiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDJUMDk6NTM6MDUuMDAwMDAwWiIsImlkIjoiNTA2NzViODllMmUyZDU3ZGJhMzQzOWRiYTFlYmRkYmYiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12017721,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "iH8DyKzUnDkCJYnwsY7NqTuBtJ3AUE5Y2qsclPbh3kJIOsLWJonWlMZ8mMZrT8yuUUoMuMtJhpZhNeK6vv73Dw",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-02T22:46:13.461Z"
  },
  "cited_facts": [
    {
      "statement": "The authors describe Qwen3.8-Flash-Next, a sparse mixture-of-experts model containing 125 billion total parameters with 6 billion activated per token.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The model leads the 397-billion parameter A17B predecessor on eight of fourteen pre-training benchmarks while trailing on the remaining six by at most 2.6 points.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Qwen3.8-Next achieves this performance at one-third the activated parameters, one-third the training tokens, and roughly one-ninth the training FLOPs compared to the predecessor.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Token mixing employs a layer-wise hybrid of Gated DeltaNet and global attention, utilizing one full-attention layer in every four layers.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "A Lemmy post compiled a list of 29 research and engineering directions for large language models, each with a brief description of its purpose and potential benefits.",
      "source": "lemmy.ml",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T09:51:41.000Z",
      "publisher_count": 1,
      "sources": [
        "lemmy.ml"
      ]
    },
    {
      "statement": "Latent reasoning performs reasoning in continuous vector representations instead of discrete token sequences, allowing more information to pass between reasoning steps.",
      "source": "lemmy.ml",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T09:51:41.000Z",
      "publisher_count": 1,
      "sources": [
        "lemmy.ml"
      ]
    },
    {
      "statement": "Linear-attention and hybrid architectures replace or combine standard attention with fixed-state mechanisms to reduce inference cost and KV-cache memory.",
      "source": "lemmy.ml",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T09:51:41.000Z",
      "publisher_count": 1,
      "sources": [
        "lemmy.ml"
      ]
    },
    {
      "statement": "The post notes that linear attention is most advantageous at very long contexts, while retaining some full-attention layers can preserve exact recall, making the optimal layer mix a central design decision.",
      "source": "lemmy.ml",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T09:51:41.000Z",
      "publisher_count": 1,
      "sources": [
        "lemmy.ml"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}