{
  "story_id": "a352e46c2147363e0bd4bb03cd4fb39f",
  "desk": "drm3",
  "revision": 1,
  "published_at": "2026-09-05T09:28:37.000Z",
  "content_hash": "09bd5bb9e4ba27bfcd05e6b674b2833d5c562273c518e66d26a181832ad2b8ce",
  "hash_basis": "sha256 over `headline\\ndek\\nprose`, plus `\\n` + the canonical citations JSON when any source is placed, plus `\\n#blog` for blogs",
  "basis": {
    "headline": "llama.cpp Adds SYCL Fusion for RMS Norm and Residual Chains",
    "dek": "llama.cpp introduces SYCL kernel fusion for RMS norm and residual addition operations.",
    "prose": "The llama.cpp project has updated its SYCL backend to fuse RMS norm, multiplication, and addition operations under the GGML_SYCL_ENABLE_FUSION flag. [^1]\n\nKernelFoundry achieved an average speedup of 2.3 on KernelBench for SYCL kernels compared to prior work. [^2]\n\nThe authors introduce KernelFoundry, an evolutionary framework designed to optimize GPU kernels by understanding hardware architecture and parallel computing strategies. [^3]\n\nThe llama.cpp release b10817 introduces two new environment variables, GGML_SYCL_MEMTRACE and GGML_SYCL_MEMTRACE_STEP, for tracing SYCL device memory usage. [^4]\n\nSupported data types for the fused operations include f32, f16, f16/f32, i32, i16, and bf16, covering both broadcast and non-contiguous memory layouts. [^5]\n\nThe optimization extends to fusing ADD+ADD residual chains using the same binbcast indexing and type matrix as the standalone add() function. [^6]\n\nKernelFoundry includes a template-based parameter optimization approach to tune kernels to specific inputs and hardware configurations. [^7]\n\nThe framework employs meta-prompt evolution to co-evolve prompts with kernels, uncovering task-specific optimization strategies. [^8]",
    "cited": "[{\"statement\":\"The llama.cpp project has updated its SYCL backend to fuse RMS norm, multiplication, and addition operations under the GGML_SYCL_ENABLE_FUSION flag.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"KernelFoundry achieved an average speedup of 2.3 on KernelBench for SYCL kernels compared to prior work.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The authors introduce KernelFoundry, an evolutionary framework designed to optimize GPU kernels by understanding hardware architecture and parallel computing strategies.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The llama.cpp release b10817 introduces two new environment variables, GGML_SYCL_MEMTRACE and GGML_SYCL_MEMTRACE_STEP, for tracing SYCL device memory usage.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-05T09:28:37.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"Supported data types for the fused operations include f32, f16, f16/f32, i32, i16, and bf16, covering both broadcast and non-contiguous memory layouts.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"The optimization extends to fusing ADD+ADD residual chains using the same binbcast indexing and type matrix as the standalone add() function.\",\"source\":\"GitHub\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:32:42.000Z\",\"publisher_count\":1,\"sources\":[\"GitHub\"]},{\"statement\":\"KernelFoundry includes a template-based parameter optimization approach to tune kernels to specific inputs and hardware configurations.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]},{\"statement\":\"The framework employs meta-prompt evolution to co-evolve prompts with kernels, uncovering task-specific optimization strategies.\",\"source\":\"arXiv.org\",\"instrument\":\"News\",\"claim_key\":null,\"published_at\":\"2026-09-04T04:00:00.000Z\",\"publisher_count\":1,\"sources\":[\"arXiv.org\"]}]",
    "kind": "news"
  },
  "receipt_verify": "Ed25519 over the dot-joined string `slice_hash.cursor_from.cursor_to.view.view_version.row_count`; public_key and sig are base64url of the raw 32-byte key / 64-byte signature",
  "receipt": null,
  "receipt_note": "this revision predates receipt-keeping (before v0.37.0); the filed row lives in the record",
  "generation_chain": {
    "wire": {
      "stream": "fountain_news",
      "story_id": "ba7408ff1876e8c66d962ae26b5f4850",
      "thread_id": "59c76ceaec2c3421f4387e5757f5ffd6",
      "thread_label": "SYCL",
      "novelty": "UPDATE",
      "content_hash": "cf5a1ef3432abe164c07e7bc504b6740b2c229aa5175d2f8124095a309475871",
      "last_published_at": "2026-09-05T09:28:37.000Z",
      "read_receipt": {
        "slice_hash": "eab18e8d2481650282cd63e8b17a1f66fa88248154323457616176eae95957fe",
        "cursor_from": "eyJ0cyI6IjIwMjYtMDktMDVUMDk6MjU6MzguMDAwMDAwWiIsImlkIjoiM2UzODQ5MGM1YWVhNDA1ZThmMzNmMjY3ODNjMWQzZmYiLCJ2IjoiMSJ9",
        "cursor_to": "eyJ0cyI6IjIwMjYtMDktMDVUMTA6Mzg6MTEuMDAwMDAwWiIsImlkIjoiZWI0Njg4MzIxMThjNDFlMTIzZTQwYTQzMTYwNWJkYWMiLCJ2IjoiMSJ9",
        "view": "v_fountain_news",
        "view_version": "1",
        "row_count": 100,
        "window_days": 3,
        "bytes_scanned": 10804931,
        "credits": 8,
        "price_per_100_rows": 8,
        "sig": "ybK8Oh-QMHgh_BBK8v9SdZDYywLup9hve34yC-DWR09FZYEx-9vXktrkTpkLLyAgZLj-JVJRkrnBVH6COiuCDA",
        "public_key": "bMUigy8O0jOnBxQ4Sc-5lwhIZ8LQVAhxMbR7qESVuUE",
        "signer_path": "lakehouse/data-extract/v1",
        "alg": "Ed25519",
        "signed": true
      }
    },
    "written_at": "2026-09-05T12:31:06.159Z"
  },
  "cited_facts": [
    {
      "statement": "The llama.cpp project has updated its SYCL backend to fuse RMS norm, multiplication, and addition operations under the GGML_SYCL_ENABLE_FUSION flag.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "KernelFoundry achieved an average speedup of 2.3 on KernelBench for SYCL kernels compared to prior work.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The authors introduce KernelFoundry, an evolutionary framework designed to optimize GPU kernels by understanding hardware architecture and parallel computing strategies.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The llama.cpp release b10817 introduces two new environment variables, GGML_SYCL_MEMTRACE and GGML_SYCL_MEMTRACE_STEP, for tracing SYCL device memory usage.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-05T09:28:37.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "Supported data types for the fused operations include f32, f16, f16/f32, i32, i16, and bf16, covering both broadcast and non-contiguous memory layouts.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "The optimization extends to fusing ADD+ADD residual chains using the same binbcast indexing and type matrix as the standalone add() function.",
      "source": "GitHub",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:32:42.000Z",
      "publisher_count": 1,
      "sources": [
        "GitHub"
      ]
    },
    {
      "statement": "KernelFoundry includes a template-based parameter optimization approach to tune kernels to specific inputs and hardware configurations.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    },
    {
      "statement": "The framework employs meta-prompt evolution to co-evolve prompts with kernels, uncovering task-specific optimization strategies.",
      "source": "arXiv.org",
      "instrument": "News",
      "claim_key": null,
      "published_at": "2026-09-04T04:00:00.000Z",
      "publisher_count": 1,
      "sources": [
        "arXiv.org"
      ]
    }
  ],
  "note": "A signature proves who filed this and that it has not changed since. It never makes a claim true."
}