{
  "story_id": "66106c92557c0ec270cc2b455d4d5aa1",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-03T10:56:08.000Z",
  "content_hash": "2e3c6acd6e86693af1f62c4dc7944e0d9e09d77336181a9da979cd126d5285e3",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "CUDA-Harness Framework Enables Agentic Kernel Generation from Natural Language",
    "dek": "Researchers propose CUDA-Harness to generate and optimize CUDA kernels directly from natural language using agentic workflows.",
    "prose": "Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier. [^1]\n\nThe authors propose CUDA-Harness, a framework for harnessing agentic CUDA kernel generation and optimization from natural language. [^2]\n\nPyTorch developers identified that tests test_get_chunk_sharding_params, test_infer_sharding_spec_from_shards_metadata, test_check_overlapping, and TestCustomShardingSpec.test_custom_sharding_spec are not guarded by any accelerator check. [^3]\n\nThe proposed solution is to keep an indexable device type for the placements that are only parsed and never used to allocate. [^4]\n\nThese tests run on CPU-only machines and building their placements from DEVICE_TYPE yields torch.device(\"cpu\", 1), which triggers a debug-build-only assert in c10/core/Device.h requiring a CPU device index of -1 or 0. [^5]\n\nCUDA-Harness introduces Intermediate-Structured Generation to connect high-level semantic understanding with low-level kernel generation. [^6]\n\nTo dilute reward hacking in Text2CUDA, the framework constructs Synthesis-Based Verification to provide isolated test data and progressive validation. [^7]\n\nGuangyey tagged a commit with the hash 819a4ee for the PyTorch repository on 2025-09-03. [^8]",
    "cited": "[{\"statement\":\"Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The authors propose CUDA-Harness, a framework for harnessing agentic CUDA kernel generation and optimization from natural language.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"PyTorch developers identified that tests test_get_chunk_sharding_params, test_infer_sharding_spec_from_shards_metadata, test_check_overlapping, and TestCustomShardingSpec.test_custom_sharding_spec are not guarded by any accelerator check.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T01:56:54.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"The proposed solution is to keep an indexable device type for the placements that are only parsed and never used to allocate.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T01:56:54.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"These tests run on CPU-only machines and building their placements from DEVICE_TYPE yields torch.device(\\\"cpu\\\", 1), which triggers a debug-build-only assert in c10/core/Device.h requiring a CPU device index of -1 or 0.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T01:56:54.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"CUDA-Harness introduces Intermediate-Structured Generation to connect high-level semantic understanding with low-level kernel generation.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"To dilute reward hacking in Text2CUDA, the framework constructs Synthesis-Based Verification to provide isolated test data and progressive validation.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-02T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Guangyey tagged a commit with the hash 819a4ee for the PyTorch repository on 2025-09-03.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-03T10:56:08.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "f2644ef4e7ac6819f5cd4be52a818e1c",
      "thread_id": "6e6766813ecf343e97551dc758f86d6b",
      "thread_label": "PyTorch",
      "novelty": "UPDATE",
      "content_hash": "b0042cbdace3646a5c9df673e39998275018e001bfb8fd65e2389e6798dd78c3",
      "last_published_at": "2026-09-03T10:56:08.000Z",
      "read_receipt": {
        "slice_hash": "256dcdd60dfda8cfb8e689b27a44613fca55ab97f7151dbaf5c2c233deded06f",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDNUMTA6NDc6NTkuMDAwMDAwWiIsImlkIjoiNDQxNWY3MmVjNjI3YWFmZmYyOWNkMjU2NGE5MWIxMDAiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDNUMTE6Mzg6NTYuMDAwMDAwWiIsImlkIjoiYTVkZWJkMWExN2NhY2MyZjlkZjJkMzZhOTk0OGNhODkiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 11903793,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "n-y37kFH7D_G1v511Ny_UEmV9ikJegt_-kDvpQVgnaqt6CypMMLVCaNja31RAB87Tt_8TcM82LQAx4bSO1o2Aw",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-03T12:21:09.294Z"
  },
  "cited_facts": [
    {
      "statement": "Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The authors propose CUDA-Harness, a framework for harnessing agentic CUDA kernel generation and optimization from natural language.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "PyTorch developers identified that tests test_get_chunk_sharding_params, test_infer_sharding_spec_from_shards_metadata, test_check_overlapping, and TestCustomShardingSpec.test_custom_sharding_spec are not guarded by any accelerator check.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T01:56:54.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "The proposed solution is to keep an indexable device type for the placements that are only parsed and never used to allocate.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T01:56:54.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "These tests run on CPU-only machines and building their placements from DEVICE_TYPE yields torch.device(\"cpu\", 1), which triggers a debug-build-only assert in c10/core/Device.h requiring a CPU device index of -1 or 0.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T01:56:54.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "CUDA-Harness introduces Intermediate-Structured Generation to connect high-level semantic understanding with low-level kernel generation.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "To dilute reward hacking in Text2CUDA, the framework constructs Synthesis-Based Verification to provide isolated test data and progressive validation.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-02T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Guangyey tagged a commit with the hash 819a4ee for the PyTorch repository on 2025-09-03.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-03T10:56:08.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}