Convert your gaming PC into an autonomous AI researcher.
This repository is a fork of jsegov/autoresearch-win-rtx (Windows), which is itself a Windows fork of the original karpathy/autoresearch. On top of the Windows port, this fork adds a self-installing WPF GUI launcher (
scripts\launch.ps1), Windows Task Scheduler integration, multi-provider AI routing viaai-powered, and tiered VRAM floors by architecture.
One day, frontier AI research used to be done by meat computers in between eating, sleeping, having other fun, and synchronizing once in a while using sound wave interconnect in the ritual of "group meeting". That era is long gone. Research is now entirely the domain of autonomous swarms of AI agents running across compute cluster megastructures in the skies. The agents claim that we are now in the 10,205th generation of the code base, in any case no one could tell if that's right or wrong as the "code" is now a self-modifying binary that has grown beyond human comprehension. This repo is the story of how it all began. -@karpathy, March 2026.
The idea: give an AI agent a small but real LLM training setup and let it experiment autonomously overnight. It modifies the code, trains for 5 minutes, checks if the result improved, keeps or discards, and repeats. You wake up in the morning to a log of experiments and (hopefully) a better model. The training code here is a simplified single-GPU implementation of nanochat. The core idea is that you're not touching any of the Python files like you normally would as a researcher. Instead, you are programming the program.md Markdown files that provide context to the AI agents and set up your autonomous research org. The default program.md in this repo is intentionally kept as a bare bones baseline, though it's obvious how one would iterate on it over time to find the "research org code" that achieves the fastest research progress, how you'd add more agents to the mix, etc. A bit more context on this project is here in this tweet and this tweet.
- Upstream source: karpathy/autoresearch (original); jsegov/autoresearch-win-rtx (direct Windows upstream).
- Primary objective: run natively on Windows with desktop consumer NVIDIA GPUs (Turing with >=8 GB VRAM, Ampere/Ada/Blackwell with >=10 GB VRAM), without unofficial Triton-on-Windows stacks.
- Scope of changes: compatibility and stability updates required for that target platform.
- The original Linux/H100-oriented path from upstream is removed in this fork and is not supported here.
- If you need the upstream Linux/H100 path, use karpathy/autoresearch.
Either command below installs this repo to %HOMEDRIVE%\myTech.Today\autoresearch-win-rtx-scheduled\ and immediately runs scripts\launch.ps1 from that location, opening the WPF GUI launcher (the easiest way to pick a provider, model, schedule, and action). Any extra CLI arguments are forwarded verbatim to the relaunched script; pass -NoGui with explicit action switches to run unattended (see How launch.ps1 Works).
-
PowerShell:
powershell -ExecutionPolicy Bypass -Command "iwr https://raw.githubusercontent.com/mytech-today-now/autoresearch-win-rtx-scheduled/refs/heads/main/scripts/launch.ps1 | iex"
-
CMD:
powershell -NoProfile -ExecutionPolicy Bypass -Command "iwr 'https://raw.githubusercontent.com/mytech-today-now/autoresearch-win-rtx-scheduled/refs/heads/main/scripts/launch.ps1' | iex"
See How launch.ps1 Works for behavior details and AI-Powered Features for provider/model configuration.
- Self-installing. On every invocation,
launch.ps1checks whether it is running from the canonical install path%HOMEDRIVE%\myTech.Today\autoresearch-win-rtx-scheduled\scripts\launch.ps1. If not (including when sourced viaiwr ... | iex), it ensuresgitis on PATH (installing viawinget install --id Git.Gitif missing), clones (orgit pull --ff-onlyupdates) the repo into the canonical path, then re-launches itself from there forwarding all CLI arguments. - Interactive WPF GUI (default). With no action switch and without
-NoGui, a WPF launcher window appears with radio buttons for the action (Preflight only, Run training now, Register scheduled task, Register task and run now, Unregister scheduled task, Update toolchain), an AI Provider group (Ollama, OpenAI, Anthropic, Azure OpenAI) and a dependent model dropdown, an editable Ollama-host combo (enabled only for Ollama), Azure endpoint/deployment fields (shown only for Azure), schedule frequency/time dropdowns (Weekly switches the time dropdown to weekday names), a maskedPasswordBoxfor API keys, an "Enable Task Scheduler history" checkbox, and OK / Cancel / Save as Defaults buttons. Save-as-Defaults persists selections (excluding the API key) to%HOMEDRIVE%\myTech.Today\autoresearch-win-rtx-scheduled\scripts\launch.json. This file is human-readable JSON and is auto-created with current values on first run if missing; both the GUI and-NoGuiCLI read it so the two surfaces behave identically (explicit CLI parameters always override stored values). - Headless / preflight modes. Pass
-NoGuiwith explicit action switches (-RunNow,-RegisterTask,-Unregister,-Update) plus provider/model/schedule parameters to run fully unattended.-NoGuiwith no action switch performs preflight only (verify tools, install if missing, ensure.venvviauv sync, start Ollama and pull the model when Ollama is selected), prints next-step hints, and exits 0.-RegisterTaskand-RunNowcan be combined to both schedule and run immediately.-Unregisterremoves the task and exits before preflight. The scheduled task always invokes the script in-NoGui -RunNowform with the chosen-Provider/-Model(and-OllamaHostfor Ollama). -Updateaction. Runsuv self update,winget upgrade --id Ollama.Ollama(Ollama provider only), andnpm install -g ai-powered@latest, refreshes the sessionPATH, then restarts the Ollama daemon and re-pulls-Modelwhen Ollama is selected. Combinable with-RegisterTaskand/or-RunNow.- Per-run logs with pruning. Each Python run writes to
%HOMEDRIVE%\myTech.Today\logs\autoresearch-run-<yyyyMMdd-HHmmss>.jsonl(local time). After each run, only the 10 newestautoresearch-run-*.jsonlfiles are retained; older ones are deleted. The aggregate append-only log%HOMEDRIVE%\myTech.Today\logs\autoresearch.jsonlis never rotated. Both surfaces capture timestampedinfo/warn/error/stdout/stderrJSON lines (UTCts). - Scheduled task is a first-class object. The task is registered via
Register-ScheduledTaskinto the visible\myTech.Today\folder asAutoresearch-Train, usingNew-ScheduledTaskPrincipal -LogonType Interactive -RunLevel Highestand only GUI-roundtrippableNew-ScheduledTaskSettingsSetoptions:-StartWhenAvailable,-AllowStartIfOnBatteries,-DontStopIfGoingOnBatteries,-RestartCount 3,-RestartInterval 5m,-MultipleInstances IgnoreNew.-ScheduleTimeis polymorphic: inHourlymode it is a minute-of-hour offset (':00'..':50'in 10-minute steps, default':00'); inDailymode it is a local 24-hour time-of-day ('00:00'..'23:45'in 15-minute steps, default'18:00'); inWeeklymode it is a weekday name ('Sunday'..'Saturday', default'Sunday') and the trigger fires that day at 03:00 local. The GUI repurposes the time dropdown and its label to match. The task is not hidden, is fully editable intaskschd.msc(no greyed-out controls), and the trigger/action/principal reflect the GUI/CLI selections. Task history is enabled by default viawevtutil set-log Microsoft-Windows-TaskScheduler/Operational /enabled:true(requires elevation; a warning is logged if elevation is unavailable).
| Parameter | Type / values | Default | Purpose |
|---|---|---|---|
-Provider |
ollama | openai | anthropic | azure |
ollama |
Selects the AI backend; exported as AI_PROVIDER (ollama maps to custom for ai-powered). |
-Model |
string | qwen2.5-coder:latest |
Model tag/name; exported as AI_POWERED_MODEL and AI_MODEL. |
-OllamaHost |
URL | http://127.0.0.1:11434 |
Ollama daemon base URL; exported as OLLAMA_HOST. |
-ApiKey |
string | (empty) | Provider key; mapped to OPENAI_API_KEY, ANTHROPIC_API_KEY, or AZURE_OPENAI_API_KEY. Never persisted. |
-AzureEndpoint / -AzureDeployment |
string | (empty) | Azure-only; exported as AZURE_OPENAI_ENDPOINT / AZURE_OPENAI_DEPLOYMENT. |
-RepoRoot |
path | parent of scripts\ |
Repo containing train.py; .venv is auto-created via uv sync if missing. |
-LogDir |
path | %HOMEDRIVE%\myTech.Today\logs |
Destination for aggregate + per-run JSONL logs. |
-ScheduleFrequency |
Hourly | Daily | Weekly |
Daily |
Trigger cadence for -RegisterTask. |
-ScheduleTime |
depends on -ScheduleFrequency (see Scheduling) |
frequency-specific (':00' / '18:00' / 'Sunday') |
Polymorphic schedule slot; validated against the active frequency at runtime. |
-RegisterTask / -RunNow / -Unregister / -Update |
switch | off | Action selectors; see above. |
-NoGui |
switch | off | Skip the WPF launcher (required for unattended/scheduled use). |
The meaning, allowed value set, and default of -ScheduleTime change with -ScheduleFrequency. The GUI's time dropdown and its label re-bind whenever the frequency selector changes; invalid CLI values raise an error listing the allowed set.
| Frequency | -ScheduleTime meaning |
Allowed values | Default |
|---|---|---|---|
Hourly |
Minute-of-the-hour offset | ':00', ':10', ':20', ':30', ':40', ':50' (10-minute steps) |
':00' |
Daily |
Local 24-hour time-of-day | '00:00', '00:15', '00:30', ... '23:45' (15-minute steps; 96 values) |
'18:00' |
Weekly |
Weekday name | 'Sunday', 'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday' |
'Sunday' |
CLI examples:
pwsh -File .\scripts\launch.ps1 -NoGui -RegisterTask -ScheduleFrequency Hourly -ScheduleTime ':30'
pwsh -File .\scripts\launch.ps1 -NoGui -RegisterTask -ScheduleFrequency Daily -ScheduleTime '18:00'
pwsh -File .\scripts\launch.ps1 -NoGui -RegisterTask -ScheduleFrequency Weekly -ScheduleTime 'Wednesday'The repository routes all AI orchestration through the ai-powered npm CLI (pinned in package.json and configured by ai-powered.json at the repo root). train.py is launched as uv run train.py so the model mediates orchestration via ai-powered's SDK bindings.
Supported providers and the GUI/CLI surface:
| Provider | -Provider value |
GUI model defaults | Required credentials |
|---|---|---|---|
| Ollama (local) | ollama |
qwen2.5-coder:latest, qwen2.5-coder:7b, qwen2.5-coder:3b, llama3.1:8b, llama3.2:3b |
none (uses -OllamaHost, default http://127.0.0.1:11434) |
| OpenAI | openai |
gpt-4o, gpt-4o-mini, gpt-4-turbo, o1, o1-mini |
OPENAI_API_KEY (set from -ApiKey or GUI PasswordBox) |
| Anthropic | anthropic |
claude-3-5-sonnet-latest, claude-3-5-haiku-latest, claude-3-opus-latest |
ANTHROPIC_API_KEY |
| Azure OpenAI | azure |
gpt-4o, gpt-4o-mini, gpt-4-turbo |
AZURE_OPENAI_API_KEY, plus -AzureEndpoint / -AzureDeployment |
The selected provider/model is exported to the child process as AI_PROVIDER, AI_POWERED_MODEL, and AI_MODEL. For Ollama, OLLAMA_HOST is also exported and the script will install/start the local ollama daemon and ollama pull the requested model during preflight. For non-Ollama providers the Ollama preflight steps are skipped.
API keys are never persisted to launch.json; only the provider, model, host, schedule, log directory, and Azure endpoint/deployment selections are saved. Supply the API key via the masked GUI field or the -ApiKey parameter at invocation time. See How launch.ps1 Works for full launcher behavior.
The repo is deliberately kept small and only really has three files that matter:
prepare.py— fixed constants, one-time data prep (downloads TinyStories data, trains a BPE tokenizer), and runtime utilities (dataloader, evaluation).train.py— the single file the agent edits. Contains the full GPT model, optimizer (Muon + AdamW), and training loop. Everything is fair game: architecture, hyperparameters, optimizer, batch size, etc. This file is edited and iterated on by the agent.program.md— baseline instructions for one agent. Point your agent here and let it go. This file is edited and iterated on by the human.
By design, training runs for a fixed 5-minute time budget (wall clock, excluding startup/compilation), regardless of the details of your compute. The metric is val_bpb (validation bits per byte) — lower is better, and vocab-size-independent so architectural changes are fairly compared.
If you are new to neural networks, this "Dummy's Guide" looks pretty good for a lot more context.
Requirements: A single NVIDIA GPU, Python 3.10+, uv.
- Single runtime path uses PyTorch SDPA attention and eager execution (no FA3/
torch.compilefast path). - Native Windows support targets desktop consumer GPUs with a tiered VRAM policy (Turing >=8 GB, Ampere/Ada/Blackwell >=10 GB), official PyTorch CUDA wheels, and SDPA attention.
- Default dataset is now TinyStories GPT-4 clean for practical consumer-GPU setup.
# 1. Install uv project manager (if you don't already have it)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
# 2. Install dependencies
uv sync
# 3. Download data and train tokenizer (one-time)
# Default dataset: TinyStories GPT-4 clean
uv run prepare.py
# 4. Manually run a single training experiment (~5 min)
uv run train.pyQuick validation run (recommended after setup):
uv run train.py --smoke-testIf the above commands all work ok, your setup is working and you can go into autonomous research mode.
Simply spin up your Claude/Codex or whatever you want in this repo (and disable all permissions), then you can prompt something like:
Hi have a look at program.md and let's kick off a new experiment! let's do the setup first.
The program.md file is essentially a super lightweight "skill".
prepare.py — constants, data prep + runtime utilities (do not modify)
train.py — model, optimizer, training loop (agent modifies this)
program.md — agent instructions
pyproject.toml — dependencies
- Single file to modify. The agent only touches
train.py. This keeps the scope manageable and diffs reviewable. - Fixed time budget. Training always runs for exactly 5 minutes, regardless of your specific platform. This means you can expect approx 12 experiments/hour and approx 100 experiments while you sleep. There are two upsides of this design decision. First, this makes experiments directly comparable regardless of what the agent changes (model size, batch size, architecture, etc). Second, this means that autoresearch will find the most optimal model for your platform in that time budget. The downside is that your runs (and results) become not comparable to other people running on other compute platforms.
- Self-contained. No external dependencies beyond PyTorch and a few small packages. No distributed training, no complex configs. One GPU, one file, one metric.
This fork's platform policy is explicit and tiered.
| Architecture | Minimum VRAM floor | Supported desktop consumer GPUs |
|---|---|---|
| Turing | >=8 GB |
RTX 2060 12GB, RTX 2060 SUPER 8GB, RTX 2070 8GB, RTX 2070 SUPER 8GB, RTX 2080 8GB, RTX 2080 SUPER 8GB, RTX 2080 Ti 11GB |
| Ampere | >=10 GB |
RTX 3060 12GB, RTX 3080 10GB, RTX 3080 12GB, RTX 3080 Ti 12GB, RTX 3090 24GB, RTX 3090 Ti 24GB |
| Ada | >=10 GB |
RTX 4060 Ti 16GB, RTX 4070 12GB, RTX 4070 SUPER 12GB, RTX 4070 Ti 12GB, RTX 4070 Ti SUPER 16GB, RTX 4080 16GB, RTX 4080 SUPER 16GB, RTX 4090 24GB |
| Blackwell | >=10 GB |
RTX 5060 Ti 16GB, RTX 5070 12GB, RTX 5070 Ti 16GB, RTX 5080 16GB, RTX 5090 32GB |
- Desktop only: laptop GPUs are not officially supported due to wide power and thermal variance.
- Floor policy: Turing desktop GPUs are supported at >=8 GB VRAM; Ampere/Ada/Blackwell desktop GPUs require >=10 GB VRAM.
RTX 2060 6GBremains out of matrix support due to VRAM floor.- Runtime path is intentionally unified across platforms: PyTorch SDPA attention + eager optimizer steps.
- Runtime adaptation is profile-driven: compute capability, BF16/TF32 support, OS, and VRAM tier determine candidate batch sizes and checkpointing strategy.
- Supported consumer profiles run a short eager-mode autotune pass and cache the selected candidate per GPU/runtime fingerprint.
- Autotune env controls:
AUTORESEARCH_DISABLE_AUTOTUNE=1skips probing;AUTORESEARCH_AUTOTUNE_REFRESH=1refreshes the cached decision. - Tested hardware in this repo remains RTX 3080 10 GB on Windows. Other listed SKUs are matrix-supported but may be less field-tested here.
- Non-goals for this fork include FA3/H100-specialized paths, unofficial Triton-for-Windows stacks, AMD/ROCm, Apple Metal, and multi-GPU training.
- Default dataset is
karpathy/tinystories_gpt4_cleanfor consumer-GPU practicality.
Seeing as there seems to be a lot of interest in tinkering with autoresearch on much smaller compute platforms than an H100, a few extra words. If you're going to try running autoresearch on smaller computers (Macbooks etc.), I'd recommend one of the forks below. On top of this, here are some recommendations for how to tune the defaults for much smaller models for aspiring forks:
- To get half-decent results I'd use a dataset with a lot less entropy, e.g. this TinyStories dataset. These are GPT-4 generated short stories. Because the data is a lot narrower in scope, you will see reasonable results with a lot smaller models (if you try to sample from them after training).
- You might experiment with decreasing
vocab_size, e.g. from 8192 down to 4096, 2048, 1024, or even - simply byte-level tokenizer with 256 possibly bytes after utf-8 encoding. - In
prepare.py, you'll want to lowerMAX_SEQ_LENa lot, depending on the computer even down to 256 etc. As you lowerMAX_SEQ_LEN, you may want to experiment with increasingDEVICE_BATCH_SIZEintrain.pyslightly to compensate. The number of tokens per fwd/bwd pass is the product of these two. - Also in
prepare.py, you'll want to decreaseEVAL_TOKENSso that your validation loss is evaluated on a lot less data. - In
train.py, the primary single knob that controls model complexity is theDEPTH(default 8, here). A lot of variables are just functions of this, so e.g. lower it down to e.g. 4. - You'll want to most likely use
WINDOW_PATTERNof just "L", because "SSSL" uses alternating banded attention pattern that may be very inefficient for you. Try it. - You'll want to lower
TOTAL_BATCH_SIZEa lot, but keep it powers of 2, e.g. down to2**14(~16K) or so even, hard to tell.
I think these would be the reasonable hyperparameters to play with. Ask your favorite coding agent for help and copy paste them this guide, as well as the full source code.
- miolini/autoresearch-macos (MacOS)
- trevin-creator/autoresearch-mlx (MacOS)
- jsegov/autoresearch-win-rtx (Windows)
- andyluo7/autoresearch (AMD)
MIT
