SYSTEM Cited by 2 sources
GitHub Copilot¶
Definition¶
GitHub Copilot is GitHub's AI coding assistant, originally an IDE autocomplete tool and later an agentic offering. It is a subscription-based product distinct from raw model API access. (Source: sources/2026-08-13-zalando-agentic-engineering-at-zalando-a-snapshot)
One shared harness across surfaces¶
Multiple Copilot products — Copilot CLI, the GitHub Copilot app, and Copilot code review — run on the same underlying agent harness, so harness-level efficiency changes propagate across all of them. GitHub still re-measures each change per surface, because a change that helps one workflow can raise cost in another (workload-local-evidence). (Source: sources/2026-09-02-github-how-we-make-ai-coding-more-cost-efficient)
Cost-efficiency engineering (2026-09-02)¶
GitHub shipped four independent harness changes that cut per-task AI cost with no detected quality regression, unified by one thesis: optimize the completed task, not the individual tool call (the local metric trap). Each was validated with task-level measurement — offline agentic-coding benchmarks, then online A/B experiments.
| Change | Mechanism | Effect (AI-credit metric) |
|---|---|---|
| Selective output compaction | Classify output; preserve source-like/arbitrary, reorganize search losslessly, compress build/test/lint noise; keep a recovery path | ~5.5% |
Remove view line-number prefixes |
Drop unused per-line numbering (current editors match code, not numbers) | ~3.1% (~5% offline inference) |
Compact the task-tool prompt |
Self-rewrite prompt ~50% smaller, guarded by a behavioral regression test | ~2.9% (~1,300 tokens/turn) |
| Reduce notification round-trips | Batch background shell/sub-agent completions and deliver results directly | ~2.3% |
The task tool launches specialized agents for parallel work; compressing its
prompt once serialized independent agents until a one-sentence rewrite
("Independent agents can run in parallel; consider side effects") restored
parallelism. GitHub also evaluated the external
RTK shell-output shortener and found it raised
end-to-end cost in their harness via reread/rerun recovery.
Seen in¶
- Zalando — an early Copilot user "from the early days when it offered autocomplete in the IDE." To complement it with API-based model access, Zalando built its LLM proxy in Jan 2024. Tools like opencode and pi are valued for mixing a Copilot subscription with the API proxy in one workflow.
Related¶
- sources/2026-09-02-github-how-we-make-ai-coding-more-cost-efficient — the four harness efficiency changes and the local-metric-trap thesis.
- local-metric-trap · cost-tracking-per-team · concepts/token-overhead
- selective-output-compaction · patterns/tool-surface-minimization · meta-prompting-loop-for-prompt-compression · batch-background-completion-delivery
- systems/rtk-rust-token-killer · systems/opencode · systems/pi-coding-agent · systems/zalando-llm-proxy
- companies/github · companies/zalando