Changelog¶
rampa 0.0.1a3 (unreleased)¶
Notes on the upcoming release will go here.
rampa 0.0.1a2 (2026-09-06)¶
rampa 0.0.1a2 moves the MCP server to FastMCP 4 and makes its tools describe
themselves: every tool declares whether calling it changes state, and
rampa://runs/{run_id} completes run ids from the live registry. The
deprecated HTTP+SSE transport is refused rather than served. Async context
managers keep their subclass type, and remote sample streaming surfaces a real
failure instead of retrying past it.
Breaking changes¶
The deprecated HTTP+SSE transport is refused¶
rampa-mcp now exits with an error when FASTMCP_TRANSPORT is set to sse
rather than serving that transport. The MCP specification deprecated HTTP+SSE
in revision 2025-03-26 and schedules it for removal ahead of every other
deprecated feature.
Switch to Streamable HTTP:
$ FASTMCP_TRANSPORT=http rampa-mcp
The default stdio transport is unaffected, as is streamable-http. (#14)
Dependencies¶
Minimum fastmcp>=4.0 (was >=2.0)¶
The mcp extra requires FastMCP 4 and gains the upper bound it never had — the
old floor admitted every release back to FastMCP 2. FastMCP 4 builds on MCP
Python SDK v2, which replaces the transport and typing stack underneath the
server.
Moving needed no changes to rampa. The server uses the API subset FastMCP 4
left untouched, and its run lifecycle was already poll-based — start_run
returns immediately and the caller polls get_status — so it needs nothing
from the session back-channel the sessionless protocol removes. (#19)
What’s new¶
MCP tools declare whether they change state¶
Every tool rampa-mcp registers now sets the MCP read_only_hint annotation.
start_run, stop_run, pause_run and resume_run declare that they change
state; the reporting tools declare that they do not. Clients read the hint to
decide whether a call needs confirmation before it runs.
The documentation renders each tool’s risk badge from the same annotation, so what the docs claim and what the server advertises cannot drift apart. (#19)
rampa://runs/{run_id} completes run ids¶
A run id is minted by the server, and MCP publishes a resource template’s URI
but never its parameter domain, so a caller had no way to discover one. The
run_id placeholder now completes from the live registry, filtered by what has
been typed.
This serves human-facing hosts — the MCP Inspector and VS Code send
completion/complete. Agent clients do not, so nothing changes for them. (#19)
Fixes¶
Async context managers keep the subclass type¶
HttpClient and
WebSocketSession annotated __aenter__
with their own class, so a subclass entering its own async with block was
inferred as the base class and lost everything it added. Both now return
Self. (#16)
Remote sample streaming surfaces real failures¶
A distributed worker treated every exception raised while draining its sample queue as “nothing available yet”, so a genuine failure was indistinguishable from an idle queue and the stream loop retried forever instead of raising. Only the empty-queue signal is treated as idle now. (#16)
Documentation¶
Tool badges carry both risk and topic¶
gp-sphinx 0.1.0a38 replaces the docs toolchain’s single tool
classification with axes this project declares, so each MCP tool renders
two badges rather than one: start_run reads mutating and
lifecycle, get_status reads readonly and lifecycle. One
classification per tool meant one had to win, so either the run-drivers
and the readers looked alike or the lifecycle grouping disappeared.
Grok CLI and Antigravity in the MCP install picker¶
The MCP install docs’ client picker now covers Grok CLI and Antigravity
(Google’s agy) alongside Claude, Codex, Gemini, and Cursor. Grok registers
through its own grok mcp add verb; Antigravity has no such verb, so its
mcpServers snippet is pasted into ~/.gemini/config/mcp_config.json. (#10)
Class fields describe themselves in the API reference (#13)¶
The event, config, metric, and threshold types now say what each field holds. They previously reached the rendered API reference as “Alias for field number 0” or as a bare name carrying only its type.
The MCP reference page links these types to their library reference entries rather than rendering a second copy of each.
Development¶
Tool documentation comes from the running server¶
The MCP tool pages are generated by introspecting a live server rather than a hand-written stand-in that had to be kept in step with it by hand. A tool’s parameters, description and risk badge all come from the same registration the server serves, so the pages cannot drift from the tools.
The docs build also fails on Sphinx warnings now, so a broken cross-reference stops the build rather than accumulating unnoticed. (#19)
FastMCP releases skip the dependency cooldown¶
fastmcp and its fastmcp-slim companion are exempt from the cooldown in
[tool.uv.exclude-newer-package]. FastMCP ships patches faster than the
cooldown window, so a lock refresh otherwise resolved to a release behind the
one being tested against. The metapackage pins its companion to an exact
version, so both need the exemption or the pair will not resolve. (#19)
Native-code boundary policy¶
Architecture decision records now define when and how rampa reaches for native code: a domain-agnostic boundary taxonomy (accelerator, engine, worker) plus load-testing guardrails that keep native speed from changing what a test measures. Charts the course for future Rust work without committing to any yet. (#6)
Self-measurement policy¶
Three architecture decision records define how rampa tests, benchmarks, and profiles itself: a deterministic end-to-end self-harness run against both the pure-Python and native paths; benchmarking that catches regressions by counting (function calls, allocations, connections) rather than on flaky wall-clock, and reserves latency claims for a named baseline; and one-command, dependency-free profiling that never distorts what a load test measures. Together they make the native-boundary preconditions from ADR 002 and ADR 003 enforceable rather than aspirational. (#7)
Target capabilities (clean-slate vision)¶
An architecture-decision record now states what rampa is for: a single contract that scales unchanged from a single process to a distributed fleet, honest measurement under load, multi-protocol support behind one interface, and a Python-first core accelerated by Rust only where measurement proves it earns its place. Each capability is anchored to a working example in a production load generator or Python/Rust project, with the technical design handed to a roadmap of follow-up records. (#7)
mcp_swap.py doctor and use-local --env¶
The dev config-swap helper gains a read-only doctor subcommand that reports
the effective MCP-swap environment — which server name each CLI points at (and a
mismatch when the repo is registered under a name other than the derived
default), un-reverted swaps and orphaned backups accumulating on disk, a state
entry whose backup has gone missing, and auth-overriding env vars such as
OPENAI_API_KEY. use-local also takes a repeatable --env KEY=VALUE to write
env (e.g. a scratch/isolation var) into the server entry without a manual
post-edit.
mcp_swap.py use-local --pr and safer config writes¶
use-local --pr N points every CLI at a pull request’s head via uvx instead
of the local checkout, so a branch can be driven through the agents without
checking it out. The number is validated as a pull request before anything is
written, and the resulting entry is launched once and handshaked against
initialize so a server that cannot start is caught before a config is
rewritten (--no-preflight skips it).
Writing is safer in four ways. A config reached through a symlink is updated in
place rather than replaced by a regular file, and the file’s mode is carried
over. A repeat swap keeps the first backup instead of taking a fresh one of the
already-swapped file, so revert still lands on the pre-swap config. Recovery
state records the resolved target path, and a backup that cannot be written, a
config that cannot be read, and a state file that cannot be parsed are each
reported against the CLI they belong to instead of aborting the run.
mcp_swap.py swaps opencode and pi¶
detect, status, use-local, revert, and doctor now cover two more agent
CLIs. opencode is written to $XDG_CONFIG_HOME/opencode/opencode.jsonc under an
mcp container, with the whole argv packed into one command array and env
spelled environment. pi is written to ~/.pi/agent/mcp.json; pi ships no MCP
client of its own, so detect says the entry only takes effect once the
third-party pi-mcp-adapter package is installed.
Both files are JSONC, so the swap edits them through a comment-preserving codec
rather than a JSON round-trip: comments, trailing commas, and formatting outside
the rewritten entry survive a swap and a revert. Per-CLI behavior moved onto
CLIInfo (its container key and dialect) instead of being spread across
if cli in (...) chains, so a new CLI is a registry entry rather than an edit
in five functions.
Warnings fail the test suite¶
Any warning raised during a test run is now an error. Upstream deprecations
surface in the release that introduces them instead of scrolling past in CI
output, and leaked HttpClient sessions fail the build
rather than being reported and ignored.
The suite was already clean, so this locks in the current state rather than working through a backlog. A dependency that starts warning will turn unrelated pull requests red; the remedy is a targeted ignore naming the upstream cause, not a broader exemption. (#15)
CI actions updated to current majors¶
Workflow actions moved to their current major releases: actions/checkout v7,
actions/cache v6, actions/setup-python v7, actions/upload-artifact v7,
astral-sh/setup-uv v9.0.0, and dorny/paths-filter v4. Workflow behavior is
unchanged, though setup-uv no longer prunes the uv cache, so the first run
after this repopulates it.
Lint floor moved to ruff 0.16¶
Minimum ruff>=0.16.0 (was unpinned). 0.16.0 formats Python code blocks inside
Markdown, so ruff format now covers README and the docs tree alongside src/
and tests/. (#16)
ruff’s default rule set is in force¶
Lint configuration moved from select to extend-select, which layers this
project’s linters on top of ruff’s curated default set instead of replacing it.
An explicit select had been silently opting the project out of the defaults.
Every exemption is scoped to one file and carries its reason in
pyproject.toml: the docs example suite runs exec by design, the event-log
drain writes synchronously on purpose, and four catch-all handlers isolate user
scenario functions, user check predicates, worker connections, and config
rollback. (#16)
rampa 0.0.1a1 (2026-05-27)¶
rampa 0.0.1a1 ships pluggable output backends, protocol clients for WebSocket and gRPC, distributed execution primitives, a Textual-based TUI dashboard, and a Rust-backed HDR histogram for O(1) metric aggregation. The engine gains mid-run pause/resume, live threshold evaluation with grace-period abort, and graceful signal handling.
What’s new¶
Output backend ecosystem¶
Eight output backends — console, JSON, CSV, InfluxDB, webhook,
Prometheus remote-write, OpenTelemetry OTLP, and GitHub Actions
annotations — plugged in via --output backend=destination. Multiple
backends run simultaneously in a single test. (#4)
WebSocket and gRPC protocol clients¶
ws and
grpc lazy properties give scenarios access
to WebSocket sessions and gRPC unary/streaming calls with automatic
metric emission matching k6’s ws_* and grpc_* vocabulary. (#4)
Distributed execution primitives¶
Deterministic work partitioning via
ExecutionSegment,
self-contained .rampa archive bundles for shipping tests to remote
workers, and a coordinator/worker wire protocol with JSON and MessagePack
encoding. (#4)
TUI dashboard¶
Textual-based live dashboard showing VU counts, iteration rate, HTTP
timing percentiles, and threshold status. Progressive display hierarchy:
console summary (default), --progress single-line, --tui full
dashboard. (#4)
Rust HDR histogram¶
PyO3-wrapped hdrhistogram crate replaces the Python list-based
TrendSink. Fixed ~20 KB memory and O(1) percentile queries regardless of
sample count, with automatic fallback to the Python implementation when
the extension is unavailable. (#4)
Pause/resume and live thresholds¶
Mid-run pause/resume via RunController, with
zero-overhead wait_if_paused() gate on each iteration. Thresholds
evaluate periodically during the run with configurable grace periods,
emitting LiveThresholdEvent for real-time
dashboard updates. (#4)
rampa inspect command¶
Show the fully resolved test configuration — scenarios, thresholds, executor defaults — in text or JSON without running the test. (#4)
CI integration¶
rampa.ci.compare compares two JSON result files and produces
text, markdown, or JSON delta reports. GitHub Actions composite action
runs a test, uploads the result artifact, and generates a step summary
with baseline comparison. (#4)
MCP discovery and control tools¶
discover_scenarios and inspect_config tools let AI agents inspect
test scripts before running. pause_run and resume_run enable mid-run
control from agent workflows. (#4)
unittest integration¶
RampaTestCase mixin adds
run_rampa() to any
unittest.TestCase, accepting a worker function and threshold
expressions for teams that don’t use pytest. (#4)
Documentation¶
Auto-generated reference docs for CLI, MCP, and pytest¶
CLI commands, MCP tools and resources, and pytest fixtures now have
auto-generated reference pages with typed parameter tables and safety
badges. The CLI was migrated from Click to argparse to enable
sphinx-autodoc-argparse directives; help output now includes
colorized example blocks. (#3)
Development¶
Build backend switched to maturin¶
Wheels now include the compiled Rust extension. maturin develop --uv
auto-builds during pytest via a conftest.py cold-start hook. (#4)
rampa 0.0.1a0 (2026-05-24)¶
rampa 0.0.1 ships the complete local load-testing engine. A user can write an async Python scenario function, run it from the CLI, and get trustworthy request metrics, checks, thresholds, and a readable summary with correct exit codes.
What’s new¶
Headless engine with typed events¶
The core engine is fully decoupled from CLI presentation.
Engine constructs per-run state and returns a
RunController with
wait(), stop(), snapshot(), and events() methods. Frontends
consume the same headless API without reaching into engine internals. (#1)
EventBus for multi-consumer event delivery¶
EventBus broadcasts
PhaseEvent,
SnapshotEvent, and
ThresholdEvent to any number of concurrent
subscribers.
Thread-safe publishing bridges the metric engine thread to the asyncio
event loop via publish_threadsafe(). (#1)
Six executor types¶
All k6-equivalent scheduling models: constant-vus, ramping-vus,
shared-iterations, per-vu-iterations, constant-arrival-rate,
and ramping-arrival-rate. Arrival-rate executors use open-model
scheduling with dropped_iterations accounting. (#1)
Typed metric model¶
Counter, Gauge, Rate, and Trend sink types behind
SinkProtocol.
The metric engine runs in a dedicated thread, draining samples from a
queue.SimpleQueue on a 50ms timer. Built-in metrics cover execution,
HTTP timing, checks, VU counts, and data transfer. (#1)
HTTP client with automatic metrics¶
HttpClient wraps aiohttp and auto-emits timing
metrics for every request: http_reqs, http_req_duration,
http_req_failed, data transfer counters, and per-phase timing via
aiohttp TraceConfig. (#1)
CLI with run, check, and doctor commands¶
rampa run executes test scripts with --vus, --duration,
--scenario, --out, --event-log, and --quiet options.
rampa check validates scripts without running them.
rampa doctor reports environment diagnostics. (#1)
MCP server¶
rampa-mcp entry point with tools for starting, stopping, and querying
load test runs. Process-local RunRegistry
tracks active and completed runs with metric snapshots and threshold
results. (#1)
pytest plugin¶
@pytest.mark.rampa_scenario marker and rampa_result fixture run
scenarios inside pytest tests. Registered via pytest11 entry point. (#1)
Threshold expressions¶
Expressions like p(95)<500 and rate<0.01 evaluate against metric
sinks and determine pass/fail exit codes. (#1)
--event-log JSONL output¶
rampa run --event-log <path> writes a JSONL event stream for
postmortem analysis — every phase transition, metric snapshot, and
threshold result as a JSON line. (#1)
Development¶
Benchmark scripts¶
Benchmark scripts measure scheduler precision, throughput, metric engine ingestion rate, and HTTP overhead. All produce JSON output for CI regression tracking. (#1)