{
  "story_id": "0004bfe6fa159c21c3d011cb6f8cb3cb",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-03T04:00:00.000Z",
  "content_hash": "cb7a627858c43ae20fa52c728738a2831cf69725b698d0c0706e259f7088f4c0",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "Self-Routing framework adapts LLM post-training per sample based on rollout behavior",
    "dek": "Self-Routing uses rollout correctness and confidence to route each LLM sample to GRPO, self-distillation, regularization, or skipping, improving math reasoning.",
    "prose": "Experiments on mathematical reasoning with Qwen3 and Qwen3.5 backbones show that Self-Routing consistently improves over uniform GRPO, uniform OPSD, fixed mixtures, and simpler routing baselines. [^1]\n\nThe paper's authors propose Self-Routing, a behavior-conditioned post-training framework for large language models that uses rollout correctness and confidence to decide how each sample should be optimized. [^2]\n\nOn LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, the REAL-Q method reduces end-to-end KL divergence by up to approximately 49% relative to state-of-the-art globally-guided methods, as claimed by the paper's authors. [^3]\n\nThe paper 'REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent', authored by Qian Zhang, Yaoming Li, and co-authors, proposes REAL-Q, a post-training quantization paradigm for large language models that targets an end-to-end-aligned surrogate of the global loss and refines it via dynamic block-wise gradient descent applied after every column block of 128 columns. [^4]\n\nThe author spent several months building a custom 2-bit quantization scheme for the Qwen3 model. [^5]\n\nThe throughput of the 2-bit quantized Qwen3 model was barely faster than the FP16 baseline. [^6]",
    "cited": "[{\"statement\":\"Experiments on mathematical reasoning with Qwen3 and Qwen3.5 backbones show that Self-Routing consistently improves over uniform GRPO, uniform OPSD, fixed mixtures, and simpler routing baselines.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The paper's authors propose Self-Routing, a behavior-conditioned post-training framework for large language models that uses rollout correctness and confidence to decide how each sample should be optimized.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, the REAL-Q method reduces end-to-end KL divergence by up to approximately 49% relative to state-of-the-art globally-guided methods, as claimed by the paper's authors.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The paper 'REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent', authored by Qian Zhang, Yaoming Li, and co-authors, proposes REAL-Q, a post-training quantization paradigm for large language models that targets an end-to-end-aligned surrogate of the global loss and refines it via dynamic block-wise gradient descent applied after every column block of 128 columns.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The author spent several months building a custom 2-bit quantization scheme for the Qwen3 model.\",\"source\":\"dzone.com\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T17:00:10.000Z\",\"publisher_count\":1,\"sources\":[\"dzone.com\"]},{\"statement\":\"The throughput of the 2-bit quantized Qwen3 model was barely faster than the FP16 baseline.\",\"source\":\"dzone.com\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T17:00:10.000Z\",\"publisher_count\":1,\"sources\":[\"dzone.com\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "00d7cb050eed263672187758d57b39e8",
      "thread_id": "5570dc9cb0f6d0523872eaf447cc83c5",
      "thread_label": "Qwen3",
      "novelty": "UPDATE",
      "content_hash": "0bab73f1f60c42c855e2a7f7c151629a9454abd431582a255b9328dc2ec13cba",
      "last_published_at": "2026-09-03T04:00:00.000Z",
      "read_receipt": {
        "slice_hash": "a6974819026a155e1c79c99aba73ab29d5f9d0aaa9cffb05e88a08023354a690",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDNUMDM6MzI6MTkuMDAwMDAwWiIsImlkIjoiNzMzZDYyYTFiNGQwMmJmNjYzNTk3YjhmN2JhZDBiZTIiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDNUMDQ6MDk6MDAuMDAwMDAwWiIsImlkIjoiNjQ5MjI3ZDNmNWQyMGQ3NmI1ZTE1NGNhODNlMTI0ZTMiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12568115,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "k4fkL1dHRNIJwSawY8K6wWBrTSdTIRM-buyZr56gcruiCoH5_IdJUeMzPH7d51ZTXAE0Qe6xSFoDys5UGTzsAw",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-03T06:46:26.217Z"
  },
  "cited_facts": [
    {
      "statement": "Experiments on mathematical reasoning with Qwen3 and Qwen3.5 backbones show that Self-Routing consistently improves over uniform GRPO, uniform OPSD, fixed mixtures, and simpler routing baselines.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The paper's authors propose Self-Routing, a behavior-conditioned post-training framework for large language models that uses rollout correctness and confidence to decide how each sample should be optimized.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, the REAL-Q method reduces end-to-end KL divergence by up to approximately 49% relative to state-of-the-art globally-guided methods, as claimed by the paper's authors.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The paper 'REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent', authored by Qian Zhang, Yaoming Li, and co-authors, proposes REAL-Q, a post-training quantization paradigm for large language models that targets an end-to-end-aligned surrogate of the global loss and refines it via dynamic block-wise gradient descent applied after every column block of 128 columns.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The author spent several months building a custom 2-bit quantization scheme for the Qwen3 model.",
      "source": "dzone.com",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T17:00:10.000Z",
      "publisher_count": 1,
      "sources": [
        "dzone.com"
      ]
    },
    {
      "statement": "The throughput of the 2-bit quantized Qwen3 model was barely faster than the FP16 baseline.",
      "source": "dzone.com",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T17:00:10.000Z",
      "publisher_count": 1,
      "sources": [
        "dzone.com"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}