SYSTEM Cited by 4 sources
Triton (language)¶
Triton is an open-source Python-embedded DSL (originally from OpenAI / Philippe Tillet) for authoring GPU kernels at a higher level of abstraction than CUDA/HIP — expressing tiled computations with auto-generated memory coalescing, shared-memory management, and tensor-core scheduling.
This page is the canonical wiki entry under the slug triton-lang that source pages reference as [systems/triton-lang](<./triton-lang.md>). For the language-and-extensions detail page (including Meta's TLX fork), see systems/triton-dsl.
Use at Pinterest — DCAT kernels¶
Pinterest's DCAT (Deduplicated Cross-Attention Transformer) is "implemented with custom Triton kernels for both training and serving" (Source: sources/2026-04-13-pinterest-scaling-recommendation-systems-with-request-level-deduplication). The custom kernels displace FlashAttention for ranking attention — DCAT's two-phase context/crossing split is not expressible as a stock FlashAttention call, so Pinterest wrote Triton kernels implementing the context pass (populate KV) + crossing pass (cross-attention against cached KV) directly.
Canonical wiki instance: Triton as the substrate for specialised attention architectures that diverge from standard self-attention (beyond what stock fused-attention kernels provide). Complements Meta's KernelEvolve use of Triton + TLX as the emission targets for LLM-generated kernels.
Caveats¶
- Kernel-level detail not disclosed — Pinterest names Triton as the implementation language but does not disclose tile sizes, shared-memory layouts, or the specific kernel shape for context vs crossing.
- Triton vs TLX vs CUDA vs HIP — the 2026-04-13 Pinterest post names only Triton. Whether Pinterest uses upstream Triton, a fork, or bridges to lower-level CUDA is not disclosed.
Use at Meta — GEM recsys training kernels (hand-written TLX)¶
Meta's GEM ads foundation model is trained with hand-written Triton Low-Level Extensions (TLX) kernels (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal):
- JFA v4 — upgrading to TLX enabled explicit warp specialization, TMA use, and persistent kernel scheduling, unlocking 30–100% TFLOPS improvement (JFA v4 is 40–140% faster than JFA v2).
- BlockAttention — a dedicated TLX kernel for fixed-64-token block-sparse self-attention eliminates FlashAttention overheads (online softmax correction, logsumexp HBM traffic, Di preprocessing) and fuses RoPE backward into the epilogue — +30.6% MFU over Triton block attention.
These are exactly the hand-authored TLX kernels that Meta's KernelEvolve aims to auto-synthesize; GEM is the human-expert baseline of the same kernel-authoring substrate.
Seen in¶
- 2026-04-13 Pinterest — Scaling Recommendation Systems with Request-Level Deduplication (sources/2026-04-13-pinterest-scaling-recommendation-systems-with-request-level-deduplication) — Triton as the DSL for DCAT's custom attention kernels.
- 2026-04-02 Meta — KernelEvolve (sources/2026-04-02-meta-kernelevolve-how-metas-ranking-engineer-agent-optimizes-ai-infrastructure) — Triton + TLX as LLM-synthesiser emission targets.
- 2026-08-03 Meta — GEM Training at LLM Scale (sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal) — hand-written TLX kernels (JFA v4, BlockAttention) with warp specialization + persistent scheduling.
- 2026-09-04 Databricks — Achieving Extreme Efficiency through Specialized GPU Kernel Generation (sources/2026-09-04-databricks-achieving-extreme-efficiency-through-specialized-gpu-kernel-generation) — Triton as the emission backend for Proteus's agent-generated, shape-specialized inference kernels (Qwen 3.5 122B, B200); C++ generations also attempted but hit build failures.
Related¶
- systems/jagged-flash-attention — GEM's TLX jagged-attention kernel.
- systems/generalized-dot-product-attention — GEM's non-softmax attention kernel.
- systems/triton-dsl — longer-form entry covering Meta's TLX fork.
- systems/tritonbench — Meta's benchmark harness for Triton kernels.
- systems/pinterest-dcat — Pinterest's custom Triton attention kernels.
- systems/flash-attention — the standard attention implementation DCAT's Triton kernels displace.
- systems/kernelevolve — Meta's automated Triton kernel synthesiser.
- systems/proteus — Databricks' agentic harness that emits Triton kernels specialized per runtime shape.