← Agentic Foundations
Foundations · Platforms

GitHub coding agents and repository work

Understand a coding-agent product, define reviewable tasks and check the changes before merging.

GitHub’s coding-agent experience is a product for work on repositories. It is a different category from a general development SDK that you embed in your own application.

This guide uses a fictional website change to explain task definition and assessment. Product features depend on your account, plan and organization settings.

What the product does

GitHub’s cloud agent can inspect a repository, plan and implement changes, and run tests in a cloud development environment. You can review its work and request refinements. GitHub warns that generated work can be incorrect or contain vulnerabilities. GitHub product documentation.

Use it as a candidate for repository tasks. Choosing it does not supply an application runtime for an unrelated business agent.

Give it a reviewable task

A useful task for a fictional news site might be:

Add a visible source label to each article card using the source already stored in article metadata. Preserve the article page layout. Handle an unknown source. Show the changed files and validation results.

This names the behavior, data source and important edge case. It gives the reviewer something observable to check.

Keep the first task limited. “Improve the entire website” mixes design judgment, content and code changes, making acceptance difficult.

Supply the right repository context

Document the commands for building and checking the project. Explain conventions that matter to the change, such as accessibility requirements and generated files.

Review which files, integrations and credentials are available to the environment. Realistic testing does not require production secrets when fixtures or a test service suffice.

Treat instructions inside fetched pages or test data as untrusted input. They should not change the task’s allowed scope.

Review the actual change

Inspect the diff before reading the agent’s summary as evidence. Check whether it edited unexpected files or added dependencies that the behavior does not require.

For the source-label task, check known and unknown sources, keyboard access and mobile layout. A build passing does not establish that the label is positioned correctly or readable.

Use existing project checks and tests appropriate to the behavior. Evaluate whether the new tests can detect a meaningful failure rather than simply repeating the implementation.

Keep merge and deployment decisions explicit

A proposed change is not a deployed change. A successful push does not prove the production site updated. Follow the project’s review, merge, build and deployment process, then inspect the public result when relevant.

For a larger codebase, assign someone who understands the affected behavior to review the result. An automated code review can add findings, but it does not establish that the task meets all product requirements.

If you use recurring repository tasks, define their scope and allowed outcomes separately from a one-off request.

Understand practical limits

Confirm current repository scope, execution limits and organization policy before assigning a long-running task. These are product constraints and can change. The official documentation is the appropriate reference.

Break work into tasks with clear acceptance criteria. A smaller change makes it easier to inspect the evidence and diagnose a failed result.

Task checklist

  • Define expected behavior and important edge cases.
  • Provide build commands and relevant conventions.
  • Limit data, integrations and credentials.
  • Review the diff and actual runtime behavior.
  • Keep merge and deployment decisions explicit.
  • Verify the final environment after deployment.

The evaluation guide applies the same outcome-first method to business agents and coding tasks.

Continue reading

Testing whether an agent completes the task ↗

Measure results, permissions, recovery and cost with repeatable normal and difficult cases.

Prepared with AI assistance and checked against the linked documentation. Examples and numerical limits are illustrative unless stated otherwise. These guides do not report independent product testing. Check current documentation before choosing a tool.

Search the publication

Find a story by title, topic, or keyword.

Press Escape to close