{
  "story_id": "d6e200e9969fab4cf0730e0b926370a6",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-02T04:00:00.000Z",
  "content_hash": "eddd241b548e4a0cf14f960f1c8f2850174c6ddfe857e1153cfca82ce0f7ade2",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "Bulgarian Man Held for Kidnapping Girlfriend and Trafficking Her for Bitcoin",
    "dek": "A 21-year-old man from Vidin was arrested for holding his girlfriend captive and forcing her to engage in online sex work to generate Bitcoin.",
    "prose": "A 21-year-old man from Vidin was arrested after a thief reported stealing his phone, which contained the keys to his cryptocurrency wallet. [^1]\n\nPolice from the Directorate for Combating Organized Crime (GDPO) investigated the origin of the Bitcoin and discovered the man had been holding his girlfriend captive since 2024. [^2]\n\nThe authors identified a failure mode in which aggregation itself induces reward hacking, where static projection aliases qualitatively different reward profiles into a single scalar. [^3]\n\nThe authors propose Adaptive Multi-Reward Projection (AMRP), a lightweight online method that reallocates aggregation weights using three signals: relative shortfall, reward volatility, and recent progress. [^4]\n\nThe man forced his girlfriend to perform sexual acts for an online webcam platform to generate income, earning between 2,000 and 6,000 euros per month. [^5]\n\nThe thief successfully transferred approximately 50,000 euros worth of Bitcoin from the victim's wallet using the stolen phone keys. [^6]\n\nAcross structured reasoning, citation-grounded generation, and open-ended alignment under GRPO, AMRP consistently improves reward-profile balance and downstream performance over fixed and dynamic weighting baselines. [^7]\n\nReinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. [^8]",
    "cited": "[{\"statement\":\"A 21-year-old man from Vidin was arrested after a thief reported stealing his phone, which contained the keys to his cryptocurrency wallet.\",\"source\":\"www.24chasa.bg\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T18:20:00.000Z\",\"publisher_count\":1,\"sources\":[\"www.24chasa.bg\"]},{\"statement\":\"Police from the Directorate for Combating Organized Crime (GDPO) investigated the origin of the Bitcoin and discovered the man had been holding his girlfriend captive since 2024.\",\"source\":\"www.24chasa.bg\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T18:20:00.000Z\",\"publisher_count\":1,\"sources\":[\"www.24chasa.bg\"]},{\"statement\":\"The authors identified a failure mode in which aggregation itself induces reward hacking, where static projection aliases qualitatively different reward profiles into a single scalar.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The authors propose Adaptive Multi-Reward Projection (AMRP), a lightweight online method that reallocates aggregation weights using three signals: relative shortfall, reward volatility, and recent progress.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The man forced his girlfriend to perform sexual acts for an online webcam platform to generate income, earning between 2,000 and 6,000 euros per month.\",\"source\":\"www.24chasa.bg\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T18:20:00.000Z\",\"publisher_count\":1,\"sources\":[\"www.24chasa.bg\"]},{\"statement\":\"The thief successfully transferred approximately 50,000 euros worth of Bitcoin from the victim's wallet using the stolen phone keys.\",\"source\":\"www.24chasa.bg\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T18:20:00.000Z\",\"publisher_count\":1,\"sources\":[\"www.24chasa.bg\"]},{\"statement\":\"Across structured reasoning, citation-grounded generation, and open-ended alignment under GRPO, AMRP consistently improves reward-profile balance and downstream performance over fixed and dynamic weighting baselines.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "55365912bbc87e521aab69e174dae09d",
      "thread_id": "581e6b9da9c3ce44f816de7a45a080e3",
      "thread_label": "GDPO",
      "novelty": "UPDATE",
      "content_hash": "7595e57f2a5ae2d254ccd76d2426b5951cb4470a72e285d33a164c324d11e7d4",
      "last_published_at": "2026-09-02T04:00:00.000Z",
      "read_receipt": {
        "slice_hash": "3905664cf10c2126ea131af76eb59a4e670141af2a4606d101f9336a8aaf6e3e",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDJUMDM6NDM6MDAuMDAwMDAwWiIsImlkIjoiZjIzZjc5YTk3YjM4MGMyNDgzZTRlNzYxMjNlZjg2M2QiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDJUMDQ6MDQ6MjYuMDAwMDAwWiIsImlkIjoiN2M3MjY3YjA5MzczMjU3ZjI0ZDZiY2NmZWY3ZDVjNGUiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12017721,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "juwMycx6aGSQfawfzLUaISZZUKKkbvDiOPDpZxKn6fmy0KIBoISR1qDjkJuogjQ0TUgCk4tYeoc9JF0bPKfOCw",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-02T22:35:58.245Z"
  },
  "cited_facts": [
    {
      "statement": "A 21-year-old man from Vidin was arrested after a thief reported stealing his phone, which contained the keys to his cryptocurrency wallet.",
      "source": "www.24chasa.bg",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T18:20:00.000Z",
      "publisher_count": 1,
      "sources": [
        "www.24chasa.bg"
      ]
    },
    {
      "statement": "Police from the Directorate for Combating Organized Crime (GDPO) investigated the origin of the Bitcoin and discovered the man had been holding his girlfriend captive since 2024.",
      "source": "www.24chasa.bg",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T18:20:00.000Z",
      "publisher_count": 1,
      "sources": [
        "www.24chasa.bg"
      ]
    },
    {
      "statement": "The authors identified a failure mode in which aggregation itself induces reward hacking, where static projection aliases qualitatively different reward profiles into a single scalar.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The authors propose Adaptive Multi-Reward Projection (AMRP), a lightweight online method that reallocates aggregation weights using three signals: relative shortfall, reward volatility, and recent progress.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The man forced his girlfriend to perform sexual acts for an online webcam platform to generate income, earning between 2,000 and 6,000 euros per month.",
      "source": "www.24chasa.bg",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T18:20:00.000Z",
      "publisher_count": 1,
      "sources": [
        "www.24chasa.bg"
      ]
    },
    {
      "statement": "The thief successfully transferred approximately 50,000 euros worth of Bitcoin from the victim's wallet using the stolen phone keys.",
      "source": "www.24chasa.bg",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T18:20:00.000Z",
      "publisher_count": 1,
      "sources": [
        "www.24chasa.bg"
      ]
    },
    {
      "statement": "Across structured reasoning, citation-grounded generation, and open-ended alignment under GRPO, AMRP consistently improves reward-profile balance and downstream performance over fixed and dynamic weighting baselines.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}