What it is
autoresearch is a small repo that gives an AI agent a simplified single-GPU implementation of nanochat and lets it experiment autonomously, typically overnight. The agent edits train.py, trains for a fixed 5 minutes, checks whether val_bpb improved, keeps or discards the change, and repeats. The human steers the process by editing program.md, a Markdown instruction file, rather than the Python files.
Who it's for
- Researchers who want to experiment with autonomous AI-agent-driven LLM research loops
- Developers with a single NVIDIA GPU who want to iterate on a 'research org' via program.md
- People adapting the setup to smaller platforms through forks or by tuning the documented hyperparameters
Requirements
Requirements
- A single NVIDIA GPU (tested on H100)
- Python 3.10+
- uv
- An AI coding agent such as Claude or Codex to run the experiments
Setup
Install uv
Install the uv project manager if you don't already have it.
bashcurl -LsSf https://astral.sh/uv/install.sh | shInstall dependencies
Sync the project dependencies with uv.
bashuv syncDownload data and train tokenizer
One-time step, about 2 minutes.
bashuv run prepare.pyRun a single training experiment
Manually run one experiment (about 5 minutes) to verify the setup works.
bashuv run train.py
Examples
Kick off the agent
PromptHi have a look at program.md and let's kick off a new experiment! let's do the setup first.Expected output: The README's suggested prompt to give your Claude/Codex agent inside the repo (with all permissions disabled), so it reads program.md and starts.
Verify the setup manually
bashuv sync
uv run prepare.py
uv run train.pyWhat it does: Installs dependencies, prepares data and the tokenizer, and runs one ~5-minute training experiment before going into autonomous mode.
Pros & cons
Pros
- Pro:Fixed 5-minute wall-clock budget makes experiments directly comparable regardless of what the agent changes, giving roughly 12 experiments per hour
- Pro:Small scope: the agent edits only train.py, which keeps diffs reviewable
- Pro:Self-contained with no distributed training or complex configs: one GPU, one file, one metric
- Pro:val_bpb is vocab-size-independent, so architectural changes are fairly compared
Cons
- Con:Requires a single NVIDIA GPU; CPU, MPS and other platforms are not supported in this repo (only via forks)
- Con:Results are not comparable to those from other people's compute platforms because of the fixed time budget
- Con:The default program.md is only a bare-bones baseline, so getting better research progress requires iterating on it yourself
Images
