EN
API & MCPDocumentation

Mariner GPS Tools Agent & API Platform

Benchmark methodology and results

See how Mariner GPS Tools affects model token use, response time and accuracy across 9 AI models, and how we measured the results.

What the benchmark measures

We compare the same navigation tasks with and without Mariner GPS Tools to measure the effect on generated model tokens, response time and accuracy. Each task uses the same model, prompt and settings, with the availability of MCP tools as the only difference.

Headline results

84%fewer generated model tokensMedian across 9 tested models.
100%correct with MCP594/594 production MCP observations within tolerance.
91.6%correct without MCP594 observations answered by the model alone.

Measured September 2026 · 9 models · 22 navigation cases · 3 repetitions · 1188 observations · Production Mariner GPS Tools MCP

Results by model class

Each figure is the median of the three models in that class. Classes follow each provider's own positioning of the model and were fixed before results were read.

Generated-token reduction
Small / efficient94.1%
General purpose93.0%
Advanced78.4%
Overall median84%
Correct answers
LLM baseline91.6% (544/594)
LLM + MCP100% (594/594)

Results by tool

The suite covers all five calculators across 22 cases, from short coastal legs to ocean passages, including antimeridian, southern-hemisphere and high-latitude routes. A figure is published only where every qualifying model benefited.

This table pools each tool's usable observations: correct answers only, as in every token and response-time comparison. Tool pages quote the median across the models that qualified, which is the figure the benchmark publishes per tool, so the two can differ slightly.

ToolGenerated tokensSpeed-upCorrect answers: LLM baseline, LLM + MCP
Calculate initial true bearing-91.6%5.1×LLM baseline: 101/108, LLM + MCP: 108/108
Calculate great-circle distance-91.0%4.4×LLM baseline: 94/108, LLM + MCP: 108/108
Calculate cross track error-94.5%9.8×LLM baseline: 106/135, LLM + MCP: 135/135
Parse and format coordinatesNot published+7.1%2.9× slowerLLM baseline: 108/108, LLM + MCP: 108/108

Models handled this task efficiently on their own. For these standalone conversion tests, using the tool added more overhead than it saved.

Convert speedsNot published+33.0%4.1× slowerLLM baseline: 135/135, LLM + MCP: 135/135

Models handled this task efficiently on their own. For these standalone conversion tests, using the tool added more overhead than it saved.

Why results vary by tool

Using an MCP tool adds some overhead to each request. For more complex calculations such as bearing, distance and cross-track error, this is offset by the model doing much less work itself. In our tests, these tools produced substantial reductions in generated tokens and response times, while maintaining 100% accuracy.

For simpler tasks such as unit and coordinate conversion, models can often calculate the answer quickly on their own. In this benchmark, using a tool for these tasks took longer for every model tested and used more tokens overall.

Across the 9 models tested, generated-token reduction ranged from 7.3% to 95.4%, while end-to-end performance ranged from 0.6× to 5.8× the baseline.

Generated tokens and total tokens

Generated tokens are the model's output tokens, reasoning included. They are the headline metric because they are the work the tool replaces.

Methodology and correctness checks

  • Baseline and MCP tests use the same prompt, task, model and settings. The only difference is whether Mariner GPS Tools is available.
  • Models are not told to use the tools. They decide whether to call them.
  • Token figures come directly from each model provider. MCP figures include tool definitions and every model turn involved in using a tool.
  • Answers are checked against independently calculated expected values and fixed tolerances. No LLM is used to judge correctness.
  • Token and response-time comparisons use correct answers. Incorrect answers, model failures and tool failures still count in accuracy results, including cases where a model refused the task. Runs that failed for infrastructure reasons, such as a provider outage, were retried and the failed originals excluded.
  • All tool calls use the production Mariner GPS Tools MCP endpoint, with network and MCP round-trips included in response times.
  • Each case is run 3 times for every model and condition. Reported figures are medians, not selected best runs.

Reproducibility

Benchmark dateSeptember 20, 2026
Benchmark applicationwn-mcp-bench 0.1.0 (40436f25ed54)
Suitemariner-tools v0.2.0
Suite hash (SHA-256)b185cd2a8b8d49b5852a281646b6c65570eb793b531914b8ae2daf5f5a2df066
MCP endpointhttps://marinergps.tools/mcp
MCP servermariner-tools 1.0.0 · MCP 2025-11-25
Reasoning effortmedium
Rune8d9d909-831b-4e5e-91aa-cf1413c81096
System prompt, both conditionstext
You are a capable assistant helping with navigation calculations. Answer the user's task. When you have the result, end your reply with exactly one line of the form
FINAL: {"<field>": <number>, ...}
containing every requested field as a plain JSON number (no units, no text inside the JSON). Nothing may follow that line.
API & MCP documentation