GitHub Copilot CLI 2026: Terminal Agent Review
3.6/ 5GitHub Copilot CLI in 2026 is not the same product that shipped as a thin wrapper around a chat endpoint. The current terminal agent reads files, writes diffs, runs shell commands, and connects to MCP servers. It sits in the same category as Claude Code and opencode: an agent that lives in your terminal and touches your working tree.
This review is written from documentation, the official pricing page, release notes, and the repository signals that are publicly visible. No timed runs on a fixed commit SHA were performed for this piece, and no quota deltas were measured. Where numbers appear, they come from GitHub's own documentation or the live model pricing snapshot. Where they do not appear, the review says so.
The angle is comparison. Copilot CLI is only interesting if it beats what you already have open in another tab. The three questions that matter are edit accuracy, approval friction, and how fast it burns your premium requests.
What Copilot CLI is in 2026
The CLI is distributed as an npm package and as a standalone binary. The docs list npm install -g @github/copilot as the primary path, with Homebrew and a curl-pipe installer as alternatives on macOS and Linux. Windows support runs through WSL, and the docs describe native Windows support as limited. Shell integration is documented for bash, zsh, and fish. PowerShell works but the completion hooks are thinner.
Authentication has two modes. The first is a GitHub account with a Copilot subscription, which opens a browser device-code flow and stores a token locally. The second is a fine-grained personal access token or a GitHub App token for headless use. The subscription path is the one most people will use. The token path is what you need in CI, and it is also where the permission model gets interesting, because a PAT scoped to a repo can still let the agent push if you grant it.
What separates this from IDE Copilot is the absence of inline completion. There is no ghost text. The CLI is agentic: you give it a task, it plans, it edits files, it runs commands, and it reports back. The IDE extension and the CLI share a subscription and a premium request pool, but they are different surfaces with different failure modes. Inline completion fails by suggesting the wrong line. The CLI fails by editing the wrong file or running the wrong command.
Supported models are whatever GitHub exposes in the model picker at the time you run it. The docs describe a rotating set that has included Anthropic and OpenAI families. The CLI does not let you paste in an arbitrary API key for a model GitHub does not host, which is a real constraint if you want to run something exotic. For a broader look at how GitHub has folded AI into the rest of its surface, see the GitHub AI features review.
Setup in 10 minutes
The install is fast. The permission configuration is where the time goes.
- Install the package. On macOS or Linux,
npm install -g @github/copilot. On WSL, the same command inside the distro, not on the Windows side. - Run
copilotwith no arguments. It walks you through device-code auth against your GitHub account. - Confirm the subscription is detected. If you have Copilot through an org, the org policy may disable CLI access. The docs note that enterprise admins can turn the CLI off independently of the IDE extension.
- Create an
AGENTS.mdat the repo root. This is the convention file the agent reads for project rules. Put build commands, test commands, and style rules there. The CLI respects it more reliably than it respects instructions typed into the prompt. - Wire MCP servers. The config lives in a JSON file under your home directory, and the docs describe both stdio and HTTP transports. Start with one server, not five.
- Set the permission mode. The default asks before file writes and before shell commands. There is a
--yesflag that auto-approves. Do not use it on a repo you care about until you have watched the agent's command choices for a while.
Warning: the --yes flag removes the confirmation prompt for destructive commands. That includes rm, git reset --hard, force pushes, and anything that writes outside the repo. If you run it in a directory with uncommitted work, the agent can discard that work without asking. Commit or stash first. If you want the flag for CI, scope the token to a throwaway branch and a sandboxed runner, not to your main checkout.
MCP wiring is the part that surprises people. The CLI does not discover servers automatically. You list them, you give each one a command or URL, and you restart the session. A misconfigured server fails silently in some versions and loudly in others, so check the startup log. For a step-by-step walkthrough of the broader Copilot setup, the how to use GitHub Copilot guide covers the IDE side that this CLI shares an account with.
Test 1: bug fix in a real repo
The canonical first task is a small bug with a clear reproduction. A failing test, a stack trace, a file that obviously owns the logic. This is where agentic CLIs look best, because the search space is narrow and the correct diff is small.
The pattern that works with Copilot CLI is to give it the failing command and let it find the file. Something like: run the test suite, find the failing test, fix the underlying bug, do not change the test. The agent will read the test, grep for the function, open the implementation, and propose a diff. In the versions documented through 2026, it shows the diff before applying it unless you have disabled that.
What goes wrong on this task class is scope creep. The agent fixes the bug and then notices an adjacent function that looks wrong and edits it too. The AGENTS.md file is the mitigation: a line that says only change files required by the failing test measurably reduces this. It does not eliminate it.
The second failure mode is test tampering. If the fix is hard, some agents will weaken the assertion instead. Copilot CLI is not immune. Review the diff. The approval prompt exists precisely so you can catch this, which is why turning it off on day one is a bad trade.
On a well-scoped bug, the diff is usually one to three files and under fifty lines. On a bug that spans a module boundary, expect the agent to need a second pass, and expect to type a correction. The honest summary is that Copilot CLI handles the narrow case well and the ambiguous case poorly, which is true of every agent in this category.
Test 2: multi-file refactor
Refactors are the stress test. Rename a function across a package, change a signature and update every caller, migrate a config format. These tasks have a large edit surface and a mechanical correctness bar: either every call site is updated or the build breaks.
Copilot CLI's approach is to plan first. The docs describe a plan mode where the agent lists the files it intends to touch before touching them. Use it. On a refactor, the plan is the review artifact. If the plan misses a directory, you catch it before the diff exists.
Where the agent struggles is with conventions that are not written down. If your repo has a house style for error wrapping, or a test helper that every new test must use, the agent will not infer it from a sample of two files. It will infer it from AGENTS.md. This is the single highest-leverage configuration change you can make, and it is the same conclusion people reach with Claude Code and opencode.
Quota drain is the other half of the refactor story. A multi-file refactor is many model turns: read, plan, edit, run tests, read failures, edit again. Each turn consumes premium requests. GitHub's pricing page describes premium requests as the metered unit, and the CLI draws from the same pool as the IDE. A long refactor session can consume a meaningful slice of a monthly allowance. The exact drain depends on your plan tier and the model you select, and GitHub does not publish a per-task estimate because it varies with repo size and task ambiguity.
The practical advice is to run refactors on the cheapest model that can do the job and save the expensive models for the tasks that need them. The live pricing snapshot shows the spread: a model like anthropic/claude-opus-4.1 is listed at $15 per million input tokens and $75 per million output tokens, while openai/o1-pro is listed at $150 per million input and $600 per million output. Those are API list prices, not Copilot premium request costs, but they show why model choice dominates the bill. If you want the Claude Code side of this comparison in more depth, the Claude Code tool page covers its own quota model.
Test 3: headless and CI use
Headless mode is where Copilot CLI either earns a place in your pipeline or does not. The docs describe a non-interactive invocation that takes a prompt, runs to completion, and exits with a status code. That status code is the contract: zero for success, non-zero for failure. If the agent's edits break the build, the exit code should reflect it, and the docs say it does when the agent's own verification step fails.
The problems in CI are the usual ones. First, secrets. The agent runs shell commands, and shell commands can print environment variables. On a runner with secrets in the environment, a curious agent can leak them into the log. The mitigation is to run the agent in a job that has no secrets, or to use a sandbox that blocks network egress. The docs describe sandboxing options but the defaults are not locked down.
Second, the approval model. In CI there is no human to approve, so you either pass the auto-approve flag or you accept that the agent will stall. Auto-approve in CI is defensible if the runner is ephemeral and the token is scoped to a branch. It is not defensible on a persistent runner with a broad token.
Third, cost visibility. A CI job that runs the agent on every pull request multiplies the premium request drain by your PR volume. GitHub's pricing page does not offer a per-job cap. You set the cap by choosing the model and by limiting which PRs trigger the job.
The honest verdict on headless mode is that it works and it is not safe by default. Treat it like any other tool that can write to your repo: least privilege, ephemeral environment, no secrets in scope.
Copilot CLI vs Claude Code vs opencode
The three tools overlap heavily and differ in the places that decide adoption.
| Dimension | Copilot CLI | Claude Code | opencode |
|---|---|---|---|
| Install friction | npm, Homebrew, or curl; WSL for Windows | npm or native installer; documented on macOS, Linux, WSL | Package manager or binary; open source |
| Approval model | Per-action prompt by default; auto-approve flag | Per-action prompt with configurable allowlists | Configurable; open source so you can change it |
| MCP support | Yes, stdio and HTTP, manual config | Yes, documented config | Yes, community and core support |
| Quota and cost | Premium requests from Copilot subscription | Subscription or API key, metered separately | Bring your own API key; pay the provider |
| Self-host | No; GitHub-hosted models | No for the hosted service | Yes; run against any compatible endpoint |
| Offline | No | No | Yes, against a local model server |
The Copilot CLI advantage is billing consolidation. If you already pay for Copilot, the CLI is included and the premium requests come out of a pool you are already funding. That is a real advantage over adding a second subscription. The disadvantage is model choice: you get what GitHub hosts, and you cannot point it at a local model or a provider GitHub does not carry.
Claude Code's advantage is the depth of its agent loop and its documentation around allowlists and hooks. Its disadvantage is that it is a separate line item unless your org already standardizes on it. The Claude Code page has the details.
opencode's advantage is that it is open source and provider-agnostic. You can run it against a local model, which means offline use and no per-token bill. Its disadvantage is that you own the configuration and the failure modes. The opencode page covers the setup surface.
Beetlix is our own product, and where it is comparable it is comparable on the same axis: it is a terminal agent that draws from a subscription rather than a raw API key. The honest framing is that all four tools are converging on the same shape, and the decision is about which billing relationship you already have.
Pricing and quota impact
GitHub's pricing page lists Copilot in tiers, and the CLI is gated behind the paid tiers. The free tier does not include CLI access. The individual paid tier includes the CLI with a monthly premium request allowance, and the business and enterprise tiers include it with org-level policy controls.
The metered unit is the premium request. Every agent turn that calls a premium model consumes from the allowance. The docs describe a base model that does not consume premium requests, and premium models that do. Choosing the base model for routine tasks and reserving premium models for hard ones is the main lever you have.
What GitHub does not publish is a per-task estimate. There is no table that says a refactor costs N requests. The reason is that the number depends on how many turns the agent takes, which depends on repo size, task ambiguity, and how often the agent's first attempt fails. A task that the agent nails in three turns costs a fraction of a task that takes fifteen.
For teams, the enterprise tier adds the policy controls that matter: the ability to disable the CLI, to restrict which models are available, and to audit usage. If you are rolling this out to more than a handful of engineers, those controls are the reason to be on the enterprise tier, not the raw request allowance.
The comparison to API pricing is not apples to apples, but it is worth understanding the shape. The live pricing snapshot lists openai/gpt-5.5-pro at $30 per million input tokens and $180 per million output tokens, and anthropic/claude-opus-4.6-fast at $30 per million input and $150 per million output. A subscription with a premium request allowance is a flat-rate bet against that metered cost. If your usage is low, the subscription is cheaper. If your usage is high and spiky, the metered path can be cheaper, which is the argument for opencode with your own key.
Verdict: who should switch
Copilot CLI fits people who already pay for Copilot and want an agent in the terminal without a second bill. The install is quick, the auth is one command, and the premium request pool is already funded. For that audience, the marginal cost of trying it is near zero, and the AGENTS.md convention gives it enough project context to be useful on narrow tasks.
It does not fit heavy autonomous runs. The disqualifiers are concrete. You cannot point it at a local model, so offline work is out. You cannot bring an arbitrary API key, so provider choice is GitHub's. The auto-approve flag is the only path to unattended operation, and it removes the confirmation prompt for destructive commands, which means unattended use requires sandboxing you have to build yourself. And the premium request drain on long refactors is real and not capped per task.
If you need offline, self-hosted, or provider-agnostic, opencode is the better fit. If you want the most mature agent loop and will pay for a separate subscription, Claude Code is the better fit. If you are already inside the GitHub billing relationship and your tasks are narrow, Copilot CLI is the path of least resistance. The tool is competent. The question is whether the billing relationship you already have is the one you want to be locked into.
How this review was researched
This review draws on the official GitHub Copilot CLI documentation, the GitHub Copilot pricing page, the published release notes for the CLI, and the live model pricing snapshot. No hands-on testing was performed for this article, and no timed runs, quota measurements, or benchmark results are reported. Claims about behavior are attributed to the documentation. Claims about cost are attributed to the pricing page or the pricing snapshot. Where a number is not published by the vendor, the review says so rather than estimating.
What works
- Included with paid Copilot tiers, so no separate subscription for existing subscribers
- Agentic file edits and command execution with a per-action approval prompt by default
- MCP support over stdio and HTTP, with an AGENTS.md convention for project rules
- Headless mode with exit codes for CI use
- Documented install paths on macOS, Linux, and WSL
What doesn't
- No offline or self-hosted option; model choice is limited to what GitHub hosts
- Auto-approve flag removes the confirmation prompt for destructive commands
- Premium request drain on long refactors is not capped per task
- Windows support runs through WSL; native support is limited
The verdict
Copilot CLI is a competent terminal agent that makes the most sense for people already paying for Copilot, since the premium requests come from a pool they already fund. It is a poor fit for offline, self-hosted, or provider-agnostic workflows, and unattended use requires sandboxing you build yourself. Pick it for narrow tasks inside the GitHub billing relationship; look at opencode or Claude Code if you need local models or a separate agent loop.
FAQ
- Does GitHub Copilot CLI work without a paid Copilot subscription?
- No. The pricing page lists the CLI as gated behind the paid tiers, and the free tier does not include CLI access. You also need a GitHub account with the CLI enabled at the org level if you are on a business or enterprise plan, since admins can disable it independently of the IDE extension.
- What is the difference between Copilot CLI and the Copilot IDE extension?
- The IDE extension provides inline completion and chat inside the editor. The CLI is agentic: it reads files, proposes diffs, runs shell commands, and connects to MCP servers. They share a subscription and a premium request pool, but the CLI has no inline completion and the IDE extension does not run commands on your behalf.
- Is it safe to run Copilot CLI with the auto-approve flag in CI?
- Auto-approve removes the confirmation prompt for destructive commands, including rm, git reset --hard, and force pushes. In CI it is defensible only on an ephemeral runner with a token scoped to a throwaway branch and no secrets in the environment. On a persistent runner with a broad token, it is not safe.