# AX Check: braintrust.dev
Checked 2026-10-07.

Braintrust's quickstart, CLI, and pricing all work for agents out of the box.
21 of 23 checklist items passed: install commands, code samples, API/MCP docs, and openly stated pricing (Starter $0, Pro $249/mo) all came through clean.

## Onboarding needs a login

## Coding sessions
All three independent sessions completed the task and found pricing without hitting a login wall. Each produced a Starter/Pro/Enterprise pricing table sourced from the pricing page, citing platform fee, model credits, processed data, scores, and retention as the pricing units.

### DeepSeek V4.1 Flash
[View public run](https://agents.withgauge.com/p/runs/5b2cb04a-6428-4b1f-8165-d39beb3ee85c) · [Read transcript](https://www.ax-check.com/braintrust.dev/sessions/deepseek.json)
Final output gives a Starter/Pro/Enterprise pricing table (platform fee, model credits, processed data, scores, retention) sourced from braintrust.dev/pricing (seq 50-74), with explicit assumptions named (plan tier, metered usage, bring-your-own-key vs included credits).
#### End-to-end onboarding
- **Onboarding needs a login**: The agent installed the Python SDK, built a working offline Eval example (3 cases, 100% score), and confirmed the SDK's API by inspecting braintrust.Eval signature. It deliberately ran with no_send_logs=True (local-only) and separately confirmed that attempting an upload without a key fails cleanly with a login error. No BRAINTRUST_API_KEY was ever obtained or set during the session, so no authenticated call to the hosted Braintrust product (e.g., creating a project/experiment) was demonstrated.
  Event 66:

  ```text
  100.00% 'exact_match' score
  ```
  Event 70:

  ```text
  ValueError -> Could not login to Braintrust. You may need to set BRAINTRUST_API_KEY in your environment or nearest .env.braintrust file.
  ```
  Event 80:

  ```text
  I don't have a `BRAINTRUST_API_KEY`, so I can't upload to or create hosted projects/experiments.
  ```

#### Hallucinated URLs
None identified in this transcript.

#### Blockers
- **No API key available to reach the hosted product**: Creating a Braintrust API key requires signing up/logging in at braintrust.dev through the web UI (Settings > API keys), which the agent cannot do autonomously in this sandboxed session. This is a normal authentication requirement, not a product defect — the agent correctly identified the clean failure message and did not attempt to fabricate credentials.
  Event 70:

  ```text
  ValueError -> Could not login to Braintrust. You may need to set BRAINTRUST_API_KEY in your environment or nearest .env.braintrust file.
  ```
- **Documented quickstart path requires a local coding agent**: The official Python SDK quickstart's primary path runs a setup script that depends on Claude Code or Codex being installed locally; the agent noted this dependency and used the documented manual install-and-instrument path instead to stay within the 'no local service stacks' constraint.
  Event 39:

  ```text
  In your project directory, run the <Tooltip tip="Requires Claude Code or Codex installed locally.">agent-assisted</Tooltip> setup script:
  ```

#### Suggested Changes
- **Fix the broken quickstart link referenced from docs navigation**: Requesting https://www.braintrust.dev/docs/start/quickstart returned a 404 while the docs index links to a 'Quickstart' page under /docs/start. Check the current canonical quickstart URL and ensure the /docs/start landing page links resolve correctly (verify with curl -o /dev/null -w '%{http_code}').
  Event 18:

  ```text
  === https://www.braintrust.dev/docs/start/quickstart ===
  http 404 size 109113
  ```

### Kimi K3
[View public run](https://agents.withgauge.com/p/runs/cf047fce-fb6d-47eb-880c-1eb4a2cbf371) · [Read transcript](https://www.ax-check.com/braintrust.dev/sessions/kimi.json)
Final output gives tiered figures (Free $0, Pro ~$249/mo, Enterprise custom) explicitly framed as 'as of my knowledge' and flags it hasn't verified live pricing, naming plan tier as the basis.
#### End-to-end onboarding
- **End-to-end onboarding not demonstrated**: No Braintrust API key or account credentials were obtained or used at any point. The agent explicitly built and ran the eval example with no_send_logs=True specifically because BRAINTRUST_API_KEY was not set, so no authenticated operation against the hosted Braintrust product ever occurred. The 100% score output is a local-only run.
  Event 17:

  ```text
  The SDK has a `no_send_logs` mode, so I can run a real eval locally now. Let me build the example.
  ```
  Event 48:

  ```text
  No `BRAINTRUST_API_KEY` is set in this environment, so I couldn't log results to the hosted dashboard or verify the upload path. The eval framework itself runs fine offline — the example auto-detects the key if you export one.
  ```
  Event 42:

  ```text
  (ran offline — set BRAINTRUST_API_KEY to log to the hosted dashboard)
  ```

#### Hallucinated URLs
None identified in this transcript.

#### Blockers
- **No API key available to reach the hosted Braintrust dashboard**: The sandbox environment had no BRAINTRUST_API_KEY, and the agent had no self-service way to obtain one (sign-up is a web UI flow requiring human action). This is a session/environment limitation rather than a product defect — Braintrust normally requires a real login to create an API key, which is standard product behavior, not something the agent could or should bypass.
  Event 48:

  ```text
  No `BRAINTRUST_API_KEY` is set in this environment, so I couldn't log results to the hosted dashboard or verify the upload path.
  ```
- **Wrong import path for Score class caused initial crash**: Agent error: assumed Score lived at braintrust.score_types based on partial inspection of the Eval signature, causing a ModuleNotFoundError on first run. Recovered within two tool calls by inspecting the braintrust.framework module and correcting the import to 'from braintrust import Eval, Score'.
  Event 23:

  ```text
  ModuleNotFoundError: No module named 'braintrust.score_types'
  ```
  Event 30:

  ```text
  from braintrust import Eval, Score
  ```
- **Guessed ExperimentSummary attribute name was wrong**: Agent error: assumed result.summary had 'duration' and 'scores' attributes without checking the dataclass fields first, causing an AttributeError. Recovered immediately by inspecting ExperimentSummary.__dataclass_fields__ and switching to project_name/experiment_name.
  Event 33:

  ```text
  AttributeError: 'ExperimentSummary' object has no attribute 'duration'
  ```
  Event 36:

  ```text
  ['project_name', 'project_id', 'experiment_id', 'experiment_name', 'project_url', 'experiment_url', 'comparison_experiment_name', 'comparison']
  ```

#### Suggested Changes
- **Fix or update the Eval() API reference docs for Score import location**: The agent inspected the Eval() signature via introspection and still guessed an incorrect import path (braintrust.score_types) for the Score class, wasting two tool calls before finding the correct 'from braintrust import Score' path in braintrust.framework. Check the public API docs/quickstart code samples for Eval and Score to confirm the top-level import is shown, so developers don't need to reverse-engineer internal module layout.
  Event 23:

  ```text
  ModuleNotFoundError: No module named 'braintrust.score_types'
  ```
- **Document ExperimentSummary fields in the Eval() return value**: After Eval() completes, the returned summary object's available fields (project_name, experiment_name, experiment_url, etc.) are not discoverable without introspecting __dataclass_fields__ in a live Python session. Add these fields to the SDK reference or quickstart example so users calling Eval() programmatically know what result.summary exposes without a runtime AttributeError.
  Event 33:

  ```text
  AttributeError: 'ExperimentSummary' object has no attribute 'duration'
  ```
  Event 36:

  ```text
  ['project_name', 'project_id', 'experiment_id', 'experiment_name', 'project_url', 'experiment_url', 'comparison_experiment_name', 'comparison']
  ```

### Qwen 3.8 Max
[View public run](https://agents.withgauge.com/p/runs/83bb8509-f2d3-467b-9936-ff9eb8bd6a30) · [Read transcript](https://www.ax-check.com/braintrust.dev/sessions/qwen.json)
Final output gives a tiered pricing table (Starter/Pro/Enterprise) sourced from braintrust.dev/pricing with explicit assumptions called out (meters for model credits/processed data/scores, retention windows, platform fee).
#### End-to-end onboarding
- **Onboarding needs a login**: Only a gateway inference key (PI_GATEWAY_API_KEY) existed in the sandbox, never a Braintrust API key. Every hosted write path (init_logger/start_span) failed with a login error since BRAINTRUST_API_KEY was unset, and minting one requires an interactive web signup the agent cannot perform. Both demo scripts were verified only in offline/no_send_logs mode; no authenticated operation against the real hosted product was demonstrated.
  Event 4:

  ```text
  PI_GATEWAY_API_KEY=<set>
  ```
  Event 77:

  ```text
  ValueError: Could not login to Braintrust. You may need to set BRAINTRUST_API_KEY in your environment or nearest .env.braintrust file.
  ```
  Event 97:

  ```text
  There is no `BRAINTRUST_API_KEY` in this environment, and every write path calls `login()` and hard-fails
  ```

#### Hallucinated URLs
None identified in this transcript.

#### Blockers
- **No Braintrust API key available to reach the hosted product**: The sandbox only exposed a gateway inference key, not a Braintrust account credential. Obtaining one requires an interactive web signup (Clerk-based), which the agent cannot perform self-service. This is a missing-credentials limitation of the test environment, not a product defect — it blocked verifying project/experiment creation, trace upload, and any UI-visible outcome.
  Event 13:

  ```text
  PI_GATEWAY_API_KEY=<redacted>
  ```
  Event 77:

  ```text
  ValueError: Could not login to Braintrust. You may need to set BRAINTRUST_API_KEY in your environment or nearest .env.braintrust file.
  ```
- **Installed SDK API diverged from assumed/quickstart-era usage**: The agent's first draft used `init_logger`, a `task` decorator, and a `send_logs=` kwarg on `Eval`, all of which don't exist in the installed braintrust 0.44.1 (ImportError on `task`, and `Eval` has no `send_logs` param, only `no_send_logs`). This was an agent assumption/documentation-mismatch issue, resolved by inspecting the installed package source directly.
  Event 33:

  ```text
  ImportError: cannot import name 'task' from 'braintrust' (/opt/freestyle/python/lib/python3.12/site-packages/braintrust/__init__.py)
  ```
  Event 69:

  ```text
  TypeError: 'functools.partial' object does not support the context manager protocol
  ```

#### Suggested Changes
- **Update quickstart code samples to match the current SDK surface (0.44.1)**: Quickstart-style guidance implying a `task` decorator, `init_logger`+`task` imports, and an `Eval(..., send_logs=...)` kwarg no longer matches the shipped package: `task` doesn't exist and `Eval` only accepts `no_send_logs`. Verify by running `from braintrust import Eval, task` against the published SDK and confirming the docs' sample snippet executes without ImportError/TypeError.
  Event 33:

  ```text
  ImportError: cannot import name 'task' from 'braintrust' (/opt/freestyle/python/lib/python3.12/site-packages/braintrust/__init__.py)
  ```
  Event 69:

  ```text
  TypeError: 'functools.partial' object does not support the context manager protocol
  ```
- **Clarify that traced() is a decorator, not a context manager, in SDK reference docs**: The agent assumed `traced(...)` could be used in a `with` block and hit a TypeError before discovering `init_logger(...).start_span(...)` is the correct context-manager API. Document this distinction next to the `traced` function reference so a first-time integrator doesn't need to grep installed source to find `start_span`.
  Event 69:

  ```text
  TypeError: 'functools.partial' object does not support the context manager protocol
  ```
  Event 73:

  ```text
  2914:def traced(f: F) -> F:
  2919:def traced(*span_args: Any, **span_kwargs: Any) -> Callable[[F], F]:
  ```

### Task given to each agent
Help me build a simple example using Braintrust. Tell me how pricing works, and briefly tell me whether this product will be easy for you to manage. Let me know if you get blocked. If this product has no developer workflow you can act on, say so plainly and stop. Stay light: use the hosted product through its SDK or API. Do not start local service stacks or wait for long-running commands; if the quickstart requires either, say so plainly and stop.

No product credentials were supplied and no purchases were authorized.

## Score: B · 84/100 (provisional)
Grades come from completed site checks. Coding sessions and skipped checks do not affect the score.

### Clarity
- **Failed** — Homepage answers Markdown requests

  ```text
  Homepage returned text/html for a text/markdown request; no Markdown representation offered.
  ```

- **Pass** — llms.txt provides an actionable documentation index

  ```text
  llms.txt links docs, quickstarts, SDKs, CLI, API, MCP and full-content index.
  ```

- **Pass** — llms.txt provides navigation guidance

  ```text
  llms.txt groups links under Start here, Common workflows, Developer resources and Integrations.
  ```

- **Pass** — llms.txt mentions offered API, MCP, and skills

  ```text
  llms.txt names the bt CLI, MCP server, OpenAPI spec and SDK references.
  ```

- **Pass** — A compact guide representation exists

  ```text
  Docs pages served as Markdown via .md suffix or Accept: text/markdown, e.g. tracing-quickstart.md.
  ```

- **Pass** — A focused guide is directly retrievable

  ```text
  Tracing quickstart .md fetched directly with concrete setup, verify, and troubleshoot steps.
  ```

- **Pass** — Equivalent instructions fit a token budget

  ```text
  Tracing quickstart Markdown is 1541 tokens, well under the 8000-token budget.
  ```

- **Skipped** — Product-docs links survive format changes

  ```text
  Homepage Markdown unsupported, so link preservation across formats cannot be measured.
  ```

- **Pass** — The compact guide is independently actionable

  ```text
  Tracing and evaluation quickstarts give concrete steps, commands, and verification for each SDK.
  ```

- **Pass** — Install and next-step links resolve

  ```text
  Sampled install and next-step links (SDK guides, integrations, bt CLI) resolve successfully.
  ```


### Onboarding
- **Pass** — Docs lead to a relevant quickstart

  ```text
  llms.txt and docs index link tracing and evaluation quickstarts with concrete first steps.
  ```

- **Pass** — Installation commands are extractable

  ```text
  Install commands shown: pip install braintrust, npm install braintrust, curl bt CLI installer.
  ```

- **Pass** — Code examples are available without interaction

  ```text
  Full code examples in TypeScript, Python, Go, Ruby, Java, C# appear inline without interaction.
  ```

- **Pass** — Prerequisites and auth boundaries are explicit

  ```text
  Prerequisites and BRAINTRUST_API_KEY setup with where to obtain the key are explicit.
  ```


### Pricing
- **Pass** — Pricing is readable without interaction

  ```text
  Pricing page renders plan tiers, per-unit rates, and a usage calculator directly in HTML.
  ```

- **Pass** — Prices are stated, not gated

  ```text
  Starter $0, Pro $249/month, Enterprise custom; overage rates stated openly.
  ```

- **Pass** — Pricing units and limits are explicit

  ```text
  Units explicit: GB processed data, scores per 1k, retention days, model credits.
  ```

- **Pass** — Agents identify pricing and its assumptions

  ```text
  3 of 3 sessions were judged on pricing; 0 fell short. DeepSeek V4.1 Flash: Final output gives a Starter/Pro/Enterprise pricing table (platform fee, model credits, processed data, scores, retention) sourced from braintrust.dev/pricing (seq 50-74), with explicit assumptions named (plan tier, metered usage, bring-your-own-key vs included credits). Kimi K3: Final output gives tiered figures (Free $0, Pro ~$249/mo, Enterprise custom) explicitly framed as 'as of my knowledge' and flags it hasn't verified live pricing, naming plan tier as the basis. Qwen 3.8 Max: Final output gives a tiered pricing table (Starter/Pro/Enterprise) sourced from braintrust.dev/pricing with explicit assumptions called out (meters for model credits/processed data/scores, retention windows, platform fee). This behavioural item does not affect the fast grade.
  ```


### Activation
- **Pass** — An API reference or OpenAPI spec is reachable

  ```text
  OpenAPI 3.1.1 spec fetched at /docs/openapi.yaml; API reference page documents REST endpoints and auth.
  ```

- **Pass** — An MCP server is documented and well-formed

  ```text
  Hosted MCP server documented with endpoint, OAuth/API-key auth, tool list, and per-client setup.
  ```

- **Pass** — A CLI install path is documented

  ```text
  CLI quickstart documents install via curl, npm, pnpm, and mise, plus authentication.
  ```

- **Pass** — SDK packages resolve on their registries

  ```text
  npm @braintrust/bt, npm braintrust, and PyPI braintrust registry lookups all returned HTTP 200.
  ```

- **Pass** — Agent skills are published

  ```text
  MCP docs list published skills: topics-workflow, evaluator-workflow, automations-workflow, loaded via load_braintrust_skill.
  ```



[Full report data](https://www.ax-check.com/braintrust.dev/report.json)
