Skip to content

GITHUB 2026-08-24 Tier 2

Read original ↗

Your alt text passes automated checks. That doesn't mean it's any good.

Summary

GitHub's accessibility team built an alt-text plugin for the GitHub Accessibility Scanner that goes beyond checking whether alt text exists to judging whether it is useful. The post is a case study in validation-pipeline design under false-positive pressure: the core design decision is drawing a hard line between checks that can prove a string is wrong (five cheap deterministic rules, on by default) and checks that can only suspect it (one vision-model rule, opt-in). Along the way it surfaces transferable engineering lessons — why the worst bug was a layout problem not a parsing problem, how to stop a model from "having opinions," why sending page data to a model is a privacy-and-cost decision, and how an offline grading harness that shares one prompt with the runtime keeps CI and local eval in sync. The tradeoffs are pitched as general to anyone "building automated checks of your own, for accessibility or otherwise."

Key takeaways

  • Separate what you can prove from what you can only suspect, and give them different defaults. Presence of alt text is an objective fact; quality is a judgment call. Five deterministic string-only rules (absent/whitespace, filename-as-alt, placeholder like TODO, single generic word like image, repeated-across-adjacent) run by default with no credentials or network. One model-backed alt-text-quality rule is opt-in. (Source: sources/2026-08-24-github-your-alt-text-passes-automated-checks-that-doesnt-mean-its-any-good)
  • A quality checker lives or dies on false positives. They chose closed sets over clever heuristics: the vague-alt rule normalizes the string then fires only on an exact match against a curated word list. alt="image" fires; alt="image of the login screen with the SSO button highlighted" does not. They knowingly accept misses to avoid false positives — "a reliable checker that developers enable beats one that gets switched off."
  • Repetition is a layout problem, not a DOM problem. The first version walked images in document order and flagged runs sharing normalized alt — but a header logo and footer logo sit adjacent in the extracted list yet nowhere near each other on screen. The rule now compares bounding-box geometry: it only extends a run when the gap between two boxes is small relative to the boxes themselves (gap > GAP_MULTIPLIER * largerDim ends the run).
  • The multiplier is a judgment call, tuned against real pages, not derived from a spec. And when either image has no measurable box the check fails open and the run continues — "a missing finding is invisible; a wrong one isn't."
  • Which images to judge is decided by the accessibility tree, not querySelectorAll('img'). They use Playwright's role-based locator, so anything not in the browser's accessibility tree drops out — including alt="", which is the author explicitly marking an image decorative. Flagging empty alt "would punish exactly the behavior you want to encourage."
  • Get the model to act like a reviewer, not a critic. Given good alt text, the naive checker always suggested "better" alt text, because "could this be better?" is a question a model always answers yes to — every image becomes a finding and the signal disappears. Three fixes: (1) a decision procedure (four ordered steps — decorative / redundant-with-caption / functional / informative — stop at first match and emit that verdict) instead of an open-ended instruction; (2) explicit anti-nitpick rules (trust the author's framing; a short alt is correct when surrounding prose already analyzes the image); (3) structured output with forced field order so reasoning is generated before verdict — the model must build an argument before it picks a label.
  • Sending images to a model is a privacy and cost decision. The rule is off by default and needs a GitHub Models token. URLs are redacted — image src/srcset and link href query+fragment are stripped (they carry signed CDN tokens / session IDs) before anything enters the model context or error logs; src/srcset become (omitted). Everything in the context window is untrusted input — page titles, headings, and prose can contain text written to steer the model, and structured output constrains the shape of a response, not the reasoning.
  • Redaction narrows what reaches the model, not what reaches your own pipeline. Findings still carry the real page URL and original HTML into the scanner's normal reporting — "you can't fix an image you can't locate." An optional Azure AI Vision OCR pre-pass sends image bytes to a second place, so a data-flow review must cover both paths.
  • Cost follows the same shape: roughly one model call per image per scan, which on an image-heavy site dominates the whole run — reason enough to put it on a schedule rather than on every commit.
  • The offline grading harness shares one prompt with the runtime rule, so "what you tune offline is what runs in CI." The harness is built from published teaching material (WebAIM, the W3C images tutorial, POET). Caveat: it tests only the model's judgment, not the whole pipeline — a case can score perfectly there and never reach the model in a real scan.

Systems / concepts / patterns extracted

Caveats (from the post)

  • Deterministic rules are literal: they catch alt that's obviously unwritten, not alt that's fluent and wrong. They read the alt attribute, not the computed accessible name, so an aria-label fix won't stop the finding.
  • The model-backed rule produces false positives — every finding is "a prompt for human attention, not a verdict."
  • Silence isn't coverage. The quality rule re-fetches images outside the browser session, so anything behind authentication can fail to load; fetch and model errors are logged and skipped, so a page can come back clean because nothing got checked.
  • Only HTML <img> tags are checked — SVG, role="img" containers, CSS backgrounds, and canvas are not covered.
  • New code with limited real-world feedback; "passing isn't conformance" — automated checks are a floor, testing with people who use assistive tech is the goal.

Source

  • companies/github
  • advisory-over-blocking — findings as suggestions, not gates
  • deterministic-filter-before-llm-reasoning — sibling idea in a triage pipeline
  • structured-checks-before-llm-reasoning — deterministic layer first, model for synthesis
Last updated · 766 distilled / 2,225 read