Docs · updated September 30, 2026
TestRuna MCP server and API
Let your coding agent close the loop: deploy, have TestRuna click through the live app in a real browser, read exactly what failed, fix it and check again.
Overview#
TestRuna works three ways, and all of them spend the same credits: the workspace on this site, an MCP server for coding agents (Claude Code, Codex, Cursor, Windsurf, VS Code), and a REST API for CI. Everything an agent or your CI does also shows up in the app's chat and test map, so you can open the workspace at any time and see the screenshots and evidence behind each result.
Quickstart#
- 1Create an API keyOpen MCP and API in the workspace and create a key. That page also fills your key into the setup commands below.
- 2Connect your agentPick your agent under Connect your agent and run the command or add the config.
- 3Ask it to testDeploy your app (a preview URL works), then tell your agent: Use TestRuna to test https://my-app.lovable.app and fix what fails. A first scan takes one to three minutes.
API keys#
Keys start with tr_ and act on your account: they can create apps, run tests and spend credits. A key is shown once, so store it like a password; we keep only a hash. You can have up to ten, name them per machine or CI system, and revoking one stops everything that uses it immediately.
Connect the MCP server#
The server speaks streamable HTTP at https://testruna.com/mcp and authenticates with your key in the Authorization header.
Run this once in a terminal. Claude Code keeps the server for every project.
claude mcp add --transport http testruna https://testruna.com/mcp \ --header "Authorization: Bearer <your API key>"
Check it with claude mcp list, or type /mcp inside Claude Code.
Tools#
| Tool | What it does | Cost |
|---|---|---|
test_app | Start testing a deployed app by URL: explore it and design its tests. Reuses the app if the URL was tested before. | 5 credits (free when reused) |
get_results | Status and the latest result of every test, with the reason for each failure. Can wait up to 50 seconds for a scan or run to finish. | Free |
run_tests | Run all tests, only the failed ones, or chosen test ids. | 1 credit per test |
get_fix_prompt | Turn failures into a precise fix request: where, what happens, what should happen, how to reproduce. | Free |
add_test | Design one more test from a plain-English sentence. | Free |
list_apps | Your apps and their ids. | Free |
get_account | Your plan and credits left. | Free |
A first scan takes one to three minutes and a run of 20 tests a couple of minutes, so agents call get_results with wait_seconds until it no longer says working. Replays of recorded tests are much faster and use no AI.
Prompts that work#
The server tells your agent how to use it, so plain requests are enough. A few that work well:
Deploy the app, then use TestRuna to test https://my-app.lovable.app. Fix every failure it finds and re-run the failed tests until they pass.
Use TestRuna to re-run the failed tests on my app, then fix what still fails.
Add a TestRuna test: searching for "MUG" in capitals still finds the Ceramic Mug. Then run it.
Authentication#
Base URL https://testruna.com, JSON in and out. Send the same key as Authorization: Bearer <key> on every request.
Endpoints#
| Method | Path | What it does |
|---|---|---|
| GET | /api/v1/me | Plan and credits. |
| GET | /api/v1/apps | Your apps, newest first. |
| POST | /api/v1/apps | Add an app and start its first scan. Body: {"url", "description"?, "allow_data_changes"?}. 5 credits. |
| GET | /api/v1/apps/{id}?wait=50 | Status and the latest result of every test. wait holds the request (max 55 s) until a scan or run finishes. |
| POST | /api/v1/apps/{id}/runs | Run tests. Body: {"tests": "all" | "failed" | ["a1.c2", …]}. 1 credit per test. |
| POST | /api/v1/apps/{id}/scan | Explore again and redesign the tests. 5 credits. |
| POST | /api/v1/apps/{id}/fix-prompt | A fix request for the current failures. Body: {"tests"?: [ids]}. Free. |
| POST | /api/v1/apps/{id}/tests | Add a test. Body: {"description"}. Free. |
Run it from CI#
After every deploy, run the tests and wait for the result. For example, in a shell step:
APP=p_AbC123 # from POST /api/v1/apps or GET /api/v1/apps
curl -X POST https://testruna.com/api/v1/apps/$APP/runs \
-H "Authorization: Bearer $TESTRUNA_KEY" -H "Content-Type: application/json" \
-d '{"tests": "all"}'
curl "https://testruna.com/api/v1/apps/$APP?wait=50" -H "Authorization: Bearer $TESTRUNA_KEY"Keep calling the second request until working is false, then fail the build if any test is failed.
GitHub Actions#
Tests every push once your host has deployed it. Add your key as the repository secret TESTRUNA_KEY and your app id as the variable TESTRUNA_APP, then commit this file. The job fails when a test fails, and the log lists each failure with its reason.
name: TestRuna
on:
push:
branches: [main]
jobs:
test:
runs-on: ubuntu-latest
steps:
- name: Wait for the deploy to go live
run: sleep 90 # or use your host's "deployment succeeded" event
- name: Run all tests and wait for the result
env:
KEY: ${{ secrets.TESTRUNA_KEY }}
APP: ${{ vars.TESTRUNA_APP }}
run: |
curl -sf -X POST "https://testruna.com/api/v1/apps/$APP/runs" \
-H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
-d '{"tests": "all"}' > /dev/null
for i in $(seq 1 30); do
curl -sf "https://testruna.com/api/v1/apps/$APP?wait=50" -H "Authorization: Bearer $KEY" > result.json
[ "$(jq -r .app.working result.json)" = "false" ] && break
done
jq -r '.tests[] | "\(.status)\t\(.title)\t\(.reason // "")"' result.json
! jq -e '.tests[] | select(.status == "failed")' result.json > /dev/nullTest statuses#
Each test in a response has an id, a title, a status and, when something went wrong, a reason.
passed | Every step ran and every expected result was confirmed on the page. |
failed | A step ran but the page showed something else. The reason quotes what it found. |
blocked | A step could not be carried out, for example because the button it names is missing, or a test login is needed. |
needs_review | Everything ran, but a result could not be confirmed either way. Look at the screenshot. |
skipped | Not run: it needs a person (CAPTCHA, Google sign-in, email), changes data while data changes are off, or would delete or pay. |
error | Something went wrong on our side. These runs are not charged. |
running / not_run | In progress, or not run in the latest run. |
Errors and limits#
- Errors come back as
{"error": "…", "code": "…"}:401bad key,402not enough credits (credits),404unknown app,409a scan or run is already going or there is nothing to run,429too many requests. - Up to 120 requests a minute per account. Scans and runs also count against the daily limits of the workspace.
- Only public addresses can be tested: our browsers cannot reach localhost or private networks, so deploy first. Test only apps you own or may test (see acceptable use).
Credits#
MCP and the API spend exactly what the workspace does: plan credits first, then extra credits. When a run needs more credits than you have, nothing starts and the response says how many are missing, with a link to top up. See pricing for plans and packs.