Docs · updated September 30, 2026

TestRuna MCP server and API

Let your coding agent close the loop: deploy, have TestRuna click through the live app in a real browser, read exactly what failed, fix it and check again.

Overview#

TestRuna works three ways, and all of them spend the same credits: the workspace on this site, an MCP server for coding agents (Claude Code, Codex, Cursor, Windsurf, VS Code), and a REST API for CI. Everything an agent or your CI does also shows up in the app's chat and test map, so you can open the workspace at any time and see the screenshots and evidence behind each result.

Quickstart#

  1. 1
    Create an API key
    Open MCP and API in the workspace and create a key. That page also fills your key into the setup commands below.
  2. 2
    Connect your agent
    Pick your agent under Connect your agent and run the command or add the config.
  3. 3
    Ask it to test
    Deploy your app (a preview URL works), then tell your agent: Use TestRuna to test https://my-app.lovable.app and fix what fails. A first scan takes one to three minutes.

API keys#

Keys start with tr_ and act on your account: they can create apps, run tests and spend credits. A key is shown once, so store it like a password; we keep only a hash. You can have up to ten, name them per machine or CI system, and revoking one stops everything that uses it immediately.

Connect the MCP server#

The server speaks streamable HTTP at https://testruna.com/mcp and authenticates with your key in the Authorization header.

Run this once in a terminal. Claude Code keeps the server for every project.

Terminal
claude mcp add --transport http testruna https://testruna.com/mcp \
  --header "Authorization: Bearer <your API key>"

Check it with claude mcp list, or type /mcp inside Claude Code.

Tools#

ToolWhat it doesCost
test_appStart testing a deployed app by URL: explore it and design its tests. Reuses the app if the URL was tested before.5 credits (free when reused)
get_resultsStatus and the latest result of every test, with the reason for each failure. Can wait up to 50 seconds for a scan or run to finish.Free
run_testsRun all tests, only the failed ones, or chosen test ids.1 credit per test
get_fix_promptTurn failures into a precise fix request: where, what happens, what should happen, how to reproduce.Free
add_testDesign one more test from a plain-English sentence.Free
list_appsYour apps and their ids.Free
get_accountYour plan and credits left.Free

A first scan takes one to three minutes and a run of 20 tests a couple of minutes, so agents call get_results with wait_seconds until it no longer says working. Replays of recorded tests are much faster and use no AI.

Prompts that work#

The server tells your agent how to use it, so plain requests are enough. A few that work well:

After a deploy
Deploy the app, then use TestRuna to test https://my-app.lovable.app.
Fix every failure it finds and re-run the failed tests until they pass.
Only what broke
Use TestRuna to re-run the failed tests on my app, then fix what still fails.
A new check
Add a TestRuna test: searching for "MUG" in capitals still finds the Ceramic Mug. Then run it.

Authentication#

Base URL https://testruna.com, JSON in and out. Send the same key as Authorization: Bearer <key> on every request.

Endpoints#

MethodPathWhat it does
GET/api/v1/mePlan and credits.
GET/api/v1/appsYour apps, newest first.
POST/api/v1/appsAdd an app and start its first scan. Body: {"url", "description"?, "allow_data_changes"?}. 5 credits.
GET/api/v1/apps/{id}?wait=50Status and the latest result of every test. wait holds the request (max 55 s) until a scan or run finishes.
POST/api/v1/apps/{id}/runsRun tests. Body: {"tests": "all" | "failed" | ["a1.c2", …]}. 1 credit per test.
POST/api/v1/apps/{id}/scanExplore again and redesign the tests. 5 credits.
POST/api/v1/apps/{id}/fix-promptA fix request for the current failures. Body: {"tests"?: [ids]}. Free.
POST/api/v1/apps/{id}/testsAdd a test. Body: {"description"}. Free.

Run it from CI#

After every deploy, run the tests and wait for the result. For example, in a shell step:

Shell
APP=p_AbC123   # from POST /api/v1/apps or GET /api/v1/apps
curl -X POST https://testruna.com/api/v1/apps/$APP/runs \
  -H "Authorization: Bearer $TESTRUNA_KEY" -H "Content-Type: application/json" \
  -d '{"tests": "all"}'
curl "https://testruna.com/api/v1/apps/$APP?wait=50" -H "Authorization: Bearer $TESTRUNA_KEY"

Keep calling the second request until working is false, then fail the build if any test is failed.

GitHub Actions#

Tests every push once your host has deployed it. Add your key as the repository secret TESTRUNA_KEY and your app id as the variable TESTRUNA_APP, then commit this file. The job fails when a test fails, and the log lists each failure with its reason.

.github/workflows/testruna.yml
name: TestRuna
on:
  push:
    branches: [main]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - name: Wait for the deploy to go live
        run: sleep 90   # or use your host's "deployment succeeded" event
      - name: Run all tests and wait for the result
        env:
          KEY: ${{ secrets.TESTRUNA_KEY }}
          APP: ${{ vars.TESTRUNA_APP }}
        run: |
          curl -sf -X POST "https://testruna.com/api/v1/apps/$APP/runs" \
            -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \
            -d '{"tests": "all"}' > /dev/null
          for i in $(seq 1 30); do
            curl -sf "https://testruna.com/api/v1/apps/$APP?wait=50" -H "Authorization: Bearer $KEY" > result.json
            [ "$(jq -r .app.working result.json)" = "false" ] && break
          done
          jq -r '.tests[] | "\(.status)\t\(.title)\t\(.reason // "")"' result.json
          ! jq -e '.tests[] | select(.status == "failed")' result.json > /dev/null

Test statuses#

Each test in a response has an id, a title, a status and, when something went wrong, a reason.

passedEvery step ran and every expected result was confirmed on the page.
failedA step ran but the page showed something else. The reason quotes what it found.
blockedA step could not be carried out, for example because the button it names is missing, or a test login is needed.
needs_reviewEverything ran, but a result could not be confirmed either way. Look at the screenshot.
skippedNot run: it needs a person (CAPTCHA, Google sign-in, email), changes data while data changes are off, or would delete or pay.
errorSomething went wrong on our side. These runs are not charged.
running / not_runIn progress, or not run in the latest run.

Errors and limits#

  • Errors come back as {"error": "…", "code": "…"}: 401 bad key, 402 not enough credits (credits), 404 unknown app, 409 a scan or run is already going or there is nothing to run, 429 too many requests.
  • Up to 120 requests a minute per account. Scans and runs also count against the daily limits of the workspace.
  • Only public addresses can be tested: our browsers cannot reach localhost or private networks, so deploy first. Test only apps you own or may test (see acceptable use).

Credits#

MCP and the API spend exactly what the workspace does: plan credits first, then extra credits. When a run needs more credits than you have, nothing starts and the response says how many are missing, with a link to top up. See pricing for plans and packs.