{
  "story_id": "30010496c5dc49267a7a99e53db43748",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-02T11:16:40.000Z",
  "content_hash": "7636ccd0393f4395c91614293341f3079bac4024f899cd2c6d99835890cf42b7",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "HybridEmo Framework Achieves Multi-Emotion Control in TTS Systems",
    "dek": "Researchers introduce HybridEmo to improve multi-emotion modeling in text-to-speech via hybrid reward optimization.",
    "prose": "The authors introduce HybridEmo, a post-training framework that initializes both tasks with supervised fine-tuning and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward. [^1]\n\nThe fix modifies the mtmd loader to keep the Qwen3-TTS code predictor ffn_down layer in F32 precision. [^2]\n\nThe llama.cpp project released a fix for the Qwen3-TTS model, specifically addressing a floating-point overflow issue in the code predictor. [^3]\n\nHuman evaluation preferred HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS. [^4]\n\nResearchers identified that emotional text-to-speech systems primarily model a single utterance-level affect, leaving multi-emotion control underexplored. [^5]\n\nSupervised fine-tuning does not explicitly evaluate emotion features, creating a supervision mismatch where single-emotion rewards provide neither structure-aware feedback for trajectory completion nor pair-aware feedback for blending. [^6]\n\nUsing F16 for the ffn_down layer caused the input to be cast to the weight type, turning the peak activation into infinity, which resulted in NaN values in the subsequent rms_norm layer. [^7]",
    "cited": "[{\"statement\":\"The authors introduce HybridEmo, a post-training framework that initializes both tasks with supervised fine-tuning and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The fix modifies the mtmd loader to keep the Qwen3-TTS code predictor ffn_down layer in F32 precision.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T11:16:40.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"The llama.cpp project released a fix for the Qwen3-TTS model, specifically addressing a floating-point overflow issue in the code predictor.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T11:16:40.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"Human evaluation preferred HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Researchers identified that emotional text-to-speech systems primarily model a single utterance-level affect, leaving multi-emotion control underexplored.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Supervised fine-tuning does not explicitly evaluate emotion features, creating a supervision mismatch where single-emotion rewards provide neither structure-aware feedback for trajectory completion nor pair-aware feedback for blending.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Using F16 for the ffn_down layer caused the input to be cast to the weight type, turning the peak activation into infinity, which resulted in NaN values in the subsequent rms_norm layer.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T11:16:40.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "a54c62ec2a872343c2b665ad3c514ec7",
      "thread_id": "962329896bd21820b0aa7d1a516fe79a",
      "thread_label": "Qwen3-TTS",
      "novelty": "UPDATE",
      "content_hash": "d67c0e9e7e556b555d873e077a7ad346b21b57a8ebfc628d907492d5f7952ca4",
      "last_published_at": "2026-09-02T11:16:40.000Z",
      "read_receipt": {
        "slice_hash": "d2caee06ef24006f0dd17c7361a34231382696021586171e4bbbf84582898f33",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDJUMTE6MDA6MjIuMDAwMDAwWiIsImlkIjoiM2RkNTFiYjNlY2NkMTFkMzc0ZmJjN2UyNWU0OTBhOWYiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDJUMTI6MDc6NDUuMDAwMDAwWiIsImlkIjoiNzY2YWRiYjk1OTBjMDc4M2RhNGEwZGY0YTM5OTFhNGMiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 96,
        "window_days": 3,
        "bytes_scanned": 12017721,
        "credits": 7.68,
        "price_per_100_rows": 8,
        "sig": "nGS7Ola_o177GOTTq3Zrjbgx0igmfv9NR6ouRT77rxPv63knBfy9LLzT53vMyeABE2KSWXeJUHm4RvGMXZJpAQ",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-02T22:46:35.296Z"
  },
  "cited_facts": [
    {
      "statement": "The authors introduce HybridEmo, a post-training framework that initializes both tasks with supervised fine-tuning and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The fix modifies the mtmd loader to keep the Qwen3-TTS code predictor ffn_down layer in F32 precision.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T11:16:40.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "The llama.cpp project released a fix for the Qwen3-TTS model, specifically addressing a floating-point overflow issue in the code predictor.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T11:16:40.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "Human evaluation preferred HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Researchers identified that emotional text-to-speech systems primarily model a single utterance-level affect, leaving multi-emotion control underexplored.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Supervised fine-tuning does not explicitly evaluate emotion features, creating a supervision mismatch where single-emotion rewards provide neither structure-aware feedback for trajectory completion nor pair-aware feedback for blending.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Using F16 for the ffn_down layer caused the input to be cast to the weight type, turning the peak activation into infinity, which resulted in NaN values in the subsequent rms_norm layer.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T11:16:40.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}