Instant Agent
A ready-to-use Colab launcher for open-weight LLMs. One command sets up vLLM and an agent on a Colab machine.
- TYPE
- Personal Tool
- YEAR
- 2026
- ROLE
- Solo
- STACK
- sh · Colab · vLLM
What it's for
You might want to run an open-weight LLM of your own on a local machine, but even the strongest consumer card, the RTX 5090, can only run smaller models at specific quantizations. Google Colab offers serious compute, such as the Pro 6000 and the A100. This script launches everything with one command and deploys an open-weight model quickly and conveniently. You can do temporary work in the cloud, or expose the API directly.
On the Colab machine the script sets up the model download, the vLLM backend, the install of opencode, oh my pi or codex, and the automatic start of herdr. The default deployment is Qwen 3.8 27B uncensored in BF16. On the A100 High-RAM configuration it measures 26 tok/s on a single stream and 232 tok/s in aggregate across 12 concurrent streams.
Why herdr
Colab’s new CLI supports SSH, but only one connection at a time; herdr’s multiple panes fill that gap. The CLI also keeps the session alive on its own, and once a task is done you can open another shell pane to copy the resulting files back.
Who it’s for
If your computer can’t run a model this size but you want to try one, if the tasks you have in mind keep triggering refusals, if you have a one-off job you’d rather not run locally, or if you simply find a locally hosted model too loud, this script works out of the box and fits well.
How to use it
- 01
Install the Colab CLI
Install uv, then install the Colab CLI and log in.
- 02
Launch
Run start.sh.
- 03
Pick a configuration
In the interactive prompts of start.sh, choose the target GPU configuration and the one-click launch type.