Mariner GPS Tools Agent & API Platform
Benchmark methodology and results
See how Mariner GPS Tools affects model token use, response time and accuracy across 9 AI models, and how we measured the results.
What the benchmark measures
We compare the same navigation tasks with and without Mariner GPS Tools to measure the effect on generated model tokens, response time and accuracy. Each task uses the same model, prompt and settings, with the availability of MCP tools as the only difference.
Headline results
Measured September 2026 · 9 models · 22 navigation cases · 3 repetitions · 1188 observations · Production Mariner GPS Tools MCP
Results by model class
Each figure is the median of the three models in that class. Classes follow each provider's own positioning of the model and were fixed before results were read.
Results by tool
The suite covers all five calculators across 22 cases, from short coastal legs to ocean passages, including antimeridian, southern-hemisphere and high-latitude routes. A figure is published only where every qualifying model benefited.
This table pools each tool's usable observations: correct answers only, as in every token and response-time comparison. Tool pages quote the median across the models that qualified, which is the figure the benchmark publishes per tool, so the two can differ slightly.
Models handled this task efficiently on their own. For these standalone conversion tests, using the tool added more overhead than it saved.
Models handled this task efficiently on their own. For these standalone conversion tests, using the tool added more overhead than it saved.
Why results vary by tool
Using an MCP tool adds some overhead to each request. For more complex calculations such as bearing, distance and cross-track error, this is offset by the model doing much less work itself. In our tests, these tools produced substantial reductions in generated tokens and response times, while maintaining 100% accuracy.
For simpler tasks such as unit and coordinate conversion, models can often calculate the answer quickly on their own. In this benchmark, using a tool for these tasks took longer for every model tested and used more tokens overall.
Across the 9 models tested, generated-token reduction ranged from 7.3% to 95.4%, while end-to-end performance ranged from 0.6× to 5.8× the baseline.
Generated tokens and total tokens
Generated tokens are the model's output tokens, reasoning included. They are the headline metric because they are the work the tool replaces.
Methodology and correctness checks
- Baseline and MCP tests use the same prompt, task, model and settings. The only difference is whether Mariner GPS Tools is available.
- Models are not told to use the tools. They decide whether to call them.
- Token figures come directly from each model provider. MCP figures include tool definitions and every model turn involved in using a tool.
- Answers are checked against independently calculated expected values and fixed tolerances. No LLM is used to judge correctness.
- Token and response-time comparisons use correct answers. Incorrect answers, model failures and tool failures still count in accuracy results, including cases where a model refused the task. Runs that failed for infrastructure reasons, such as a provider outage, were retried and the failed originals excluded.
- All tool calls use the production Mariner GPS Tools MCP endpoint, with network and MCP round-trips included in response times.
- Each case is run 3 times for every model and condition. Reported figures are medians, not selected best runs.
Reproducibility
September 20, 2026wn-mcp-bench 0.1.0 (40436f25ed54)mariner-tools v0.2.0b185cd2a8b8d49b5852a281646b6c65570eb793b531914b8ae2daf5f5a2df066https://marinergps.tools/mcpmariner-tools 1.0.0 · MCP 2025-11-25mediume8d9d909-831b-4e5e-91aa-cf1413c81096You are a capable assistant helping with navigation calculations. Answer the user's task. When you have the result, end your reply with exactly one line of the form
FINAL: {"<field>": <number>, ...}
containing every requested field as a plain JSON number (no units, no text inside the JSON). Nothing may follow that line.