{
  "story_id": "aa51ce09d5aa89aa4c6e1ae498dbd6a5",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-04T04:32:42.000Z",
  "content_hash": "25d98bb3d23429a7b10456e59a7525fd79aa44a334537f0024689b30374db485",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "llama.cpp Adds SYCL Fusion for RMS Norm and Residual Chains",
    "dek": "llama.cpp introduces SYCL kernel fusion for RMS norm and residual addition operations.",
    "prose": "The llama.cpp project has updated its SYCL backend to fuse RMS norm, multiplication, and addition operations under the GGML_SYCL_ENABLE_FUSION flag. [^1]\n\nKernelFoundry achieved an average speedup of 2.3 on KernelBench for SYCL kernels compared to prior work. [^2]\n\nThe authors introduce KernelFoundry, an evolutionary framework designed to optimize GPU kernels by understanding hardware architecture and parallel computing strategies. [^3]\n\nSupported data types for the fused operations include f32, f16, f16/f32, i32, i16, and bf16, covering both broadcast and non-contiguous memory layouts. [^4]\n\nThe optimization extends to fusing ADD+ADD residual chains using the same binbcast indexing and type matrix as the standalone add() function. [^5]\n\nKernelFoundry includes a template-based parameter optimization approach to tune kernels to specific inputs and hardware configurations. [^6]\n\nKernelFoundry utilizes MAP-Elites quality diversity search with kernel-specific behavioral dimensions to sustain exploration of the GPU kernel space. [^7]\n\nUnsupported combinations of operations or data types will automatically fall back to launching two separate add() functions. [^8]",
    "cited": "[{\"statement\":\"The llama.cpp project has updated its SYCL backend to fuse RMS norm, multiplication, and addition operations under the GGML_SYCL_ENABLE_FUSION flag.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"KernelFoundry achieved an average speedup of 2.3 on KernelBench for SYCL kernels compared to prior work.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The authors introduce KernelFoundry, an evolutionary framework designed to optimize GPU kernels by understanding hardware architecture and parallel computing strategies.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Supported data types for the fused operations include f32, f16, f16/f32, i32, i16, and bf16, covering both broadcast and non-contiguous memory layouts.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"The optimization extends to fusing ADD+ADD residual chains using the same binbcast indexing and type matrix as the standalone add() function.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"KernelFoundry includes a template-based parameter optimization approach to tune kernels to specific inputs and hardware configurations.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"KernelFoundry utilizes MAP-Elites quality diversity search with kernel-specific behavioral dimensions to sustain exploration of the GPU kernel space.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"Unsupported combinations of operations or data types will automatically fall back to launching two separate add() functions.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "b4e88ed5a879e247f7c1b6f005a2e057",
      "thread_id": "59c76ceaec2c3421f4387e5757f5ffd6",
      "thread_label": "SYCL",
      "novelty": "UPDATE",
      "content_hash": "27354618b0657a969226f52c5efb1c87f79d690886632ced34ff0d63d69e8db6",
      "last_published_at": "2026-09-04T04:32:42.000Z",
      "read_receipt": {
        "slice_hash": "226f85eeb969ef4e22c236c02a6ac83a01bfea21805771a37ad265468b4f0d1b",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDRUMDQ6MDA6MDAuMDAwMDAwWiIsImlkIjoiMjdkNmMxNTc0ZTI3OThkMTZkYWY2ZDc4YWYzODMzYzYiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDRUMDQ6NDE6MTEuMDAwMDAwWiIsImlkIjoiYzI1ZDc3ZDNmMWFkMDY3MDVkMDMwZjM5NzMxYWM2YjEiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 12166856,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "aJ2O3M7WPo_Nx1rM5M-BmixwZpOjkJ_YQzXWupf3dFzyLFVqGAMLRh5EhYVvs8oAV9PfMSVGKw-zapLEln_zAg",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-04T07:02:00.324Z"
  },
  "cited_facts": [
    {
      "statement": "The llama.cpp project has updated its SYCL backend to fuse RMS norm, multiplication, and addition operations under the GGML_SYCL_ENABLE_FUSION flag.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "KernelFoundry achieved an average speedup of 2.3 on KernelBench for SYCL kernels compared to prior work.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The authors introduce KernelFoundry, an evolutionary framework designed to optimize GPU kernels by understanding hardware architecture and parallel computing strategies.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Supported data types for the fused operations include f32, f16, f16/f32, i32, i16, and bf16, covering both broadcast and non-contiguous memory layouts.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "The optimization extends to fusing ADD+ADD residual chains using the same binbcast indexing and type matrix as the standalone add() function.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "KernelFoundry includes a template-based parameter optimization approach to tune kernels to specific inputs and hardware configurations.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "KernelFoundry utilizes MAP-Elites quality diversity search with kernel-specific behavioral dimensions to sustain exploration of the GPU kernel space.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "Unsupported combinations of operations or data types will automatically fall back to launching two separate add() functions.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}