{
  "story_id": "1db8eb3a3e387cd36fa38c22f3283971",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-01T19:20:09.000Z",
  "content_hash": "b4d1c163bd02951c6b11778e6ca236b8fe69cd02ee8d9e625c145d48f069ec59",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "Google Study: Frontier Models Recall 65% of Facts via Extended Thinking",
    "dek": "Google Research shows frontier LLMs can recover up to 65% of forgotten facts by using inference-time thinking.",
    "prose": "Researchers at Google Research and Technion published a study demonstrating that frontier large language models often fail to recall facts they have already encoded in their parameters. [^1]\n\nResearchers addressed the challenge of detecting whether modern Korean poetry is human-authored or generated by Large Language Models (LLMs). [^2]\n\nExperiments showed that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts, indicating that recall rather than encoding is the primary bottleneck for factual accuracy. [^3]\n\nScaling the Gemma3 model from 1 billion to 27 billion parameters decreased encoding failures from 85% to 23%, but simultaneously increased the share of recall failures to 40% without thinking. [^4]\n\nThe study found that providing models with extra computational effort, such as inference-time thinking, successfully retrieves 40-65% of the encoded facts that models initially fail to directly recall. [^5]\n\nExpert evaluation on GPT-5.2 preferred feature-guided poems over the unconstrained baseline for generation tasks. [^6]\n\nThe classifier attained an average AUC-ROC of 83.60 in zero-shot out-of-distribution detection across seven unseen LLMs. [^7]\n\nThis performance represented an absolute gain of 7.76 AUC points and a 10.23% relative improvement over KatFishNet, the strongest baseline in the comparison. [^8]",
    "cited": "[{\"statement\":\"Researchers at Google Research and Technion published a study demonstrating that frontier large language models often fail to recall facts they have already encoded in their parameters.\",\"source\":\"venturebeat.com\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T19:20:09.000Z\",\"publisher_count\":1,\"sources\":[\"venturebeat.com\"]},{\"statement\":\"Researchers addressed the challenge of detecting whether modern Korean poetry is human-authored or generated by Large Language Models (LLMs).\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Experiments showed that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts, indicating that recall rather than encoding is the primary bottleneck for factual accuracy.\",\"source\":\"venturebeat.com\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T19:20:09.000Z\",\"publisher_count\":1,\"sources\":[\"venturebeat.com\"]},{\"statement\":\"Scaling the Gemma3 model from 1 billion to 27 billion parameters decreased encoding failures from 85% to 23%, but simultaneously increased the share of recall failures to 40% without thinking.\",\"source\":\"venturebeat.com\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T19:20:09.000Z\",\"publisher_count\":1,\"sources\":[\"venturebeat.com\"]},{\"statement\":\"The study found that providing models with extra computational effort, such as inference-time thinking, successfully retrieves 40-65% of the encoded facts that models initially fail to directly recall.\",\"source\":\"venturebeat.com\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T19:20:09.000Z\",\"publisher_count\":1,\"sources\":[\"venturebeat.com\"]},{\"statement\":\"Expert evaluation on GPT-5.2 preferred feature-guided poems over the unconstrained baseline for generation tasks.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The classifier attained an average AUC-ROC of 83.60 in zero-shot out-of-distribution detection across seven unseen LLMs.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"This performance represented an absolute gain of 7.76 AUC points and a 10.23% relative improvement over KatFishNet, the strongest baseline in the comparison.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-01T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "6c98572b1cb2d777d8870505e37cec76",
      "thread_id": "e19656bce7115796991e7c4e874f530d",
      "thread_label": "Gemini-3",
      "novelty": "UPDATE",
      "content_hash": "7ccfbbffb231df902568c52183d8694f1ed1c684e6c3a15bec37ab0ef06451da",
      "last_published_at": "2026-09-01T19:20:09.000Z",
      "read_receipt": {
        "slice_hash": "76289ed879b9905c835ccef2de145593b67c484df5704f8785b24453b37f0082",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDFUMTk6MTA6MzkuMDAwMDAwWiIsImlkIjoiMDIyYzEwN2E3ZWUwNzgxMWY1MjQ4MDc2ZDZlYmIzZjAiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDFUMTk6NDE6MzYuMDAwMDAwWiIsImlkIjoiMDM4NDgxNTEyMzViOTI1YmVlZmVkMTA0NGFiNDQxNWUiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12017721,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "FVvbpCCsUM8wy4C_W0fs1PYDL7sYVi-adH7H9iILNQ4lOOgmndHO3Hs9HrvwU8SmsMsVQ66jRg4lDCOPjQ0_AA",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-02T22:11:27.242Z"
  },
  "cited_facts": [
    {
      "statement": "Researchers at Google Research and Technion published a study demonstrating that frontier large language models often fail to recall facts they have already encoded in their parameters.",
      "source": "venturebeat.com",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T19:20:09.000Z",
      "publisher_count": 1,
      "sources": [
        "venturebeat.com"
      ]
    },
    {
      "statement": "Researchers addressed the challenge of detecting whether modern Korean poetry is human-authored or generated by Large Language Models (LLMs).",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Experiments showed that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts, indicating that recall rather than encoding is the primary bottleneck for factual accuracy.",
      "source": "venturebeat.com",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T19:20:09.000Z",
      "publisher_count": 1,
      "sources": [
        "venturebeat.com"
      ]
    },
    {
      "statement": "Scaling the Gemma3 model from 1 billion to 27 billion parameters decreased encoding failures from 85% to 23%, but simultaneously increased the share of recall failures to 40% without thinking.",
      "source": "venturebeat.com",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T19:20:09.000Z",
      "publisher_count": 1,
      "sources": [
        "venturebeat.com"
      ]
    },
    {
      "statement": "The study found that providing models with extra computational effort, such as inference-time thinking, successfully retrieves 40-65% of the encoded facts that models initially fail to directly recall.",
      "source": "venturebeat.com",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T19:20:09.000Z",
      "publisher_count": 1,
      "sources": [
        "venturebeat.com"
      ]
    },
    {
      "statement": "Expert evaluation on GPT-5.2 preferred feature-guided poems over the unconstrained baseline for generation tasks.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The classifier attained an average AUC-ROC of 83.60 in zero-shot out-of-distribution detection across seven unseen LLMs.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "This performance represented an absolute gain of 7.76 AUC points and a 10.23% relative improvement over KatFishNet, the strongest baseline in the comparison.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-01T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}