PUBLIC RESEARCH · AI CONNECTIONS

Explore the lab with your AI.

Ask about the published results, compare the settings, and follow the sources.

CONNECT YOUR ASSISTANT

Bring the evidence into your conversation.

MCP lets a compatible AI assistant look up DeepWakeLabs results and published articles while you chat. The tools read the lab’s published data; they do not run new benchmarks.

  1. In an AI client that supports remote MCP connections, add the endpoint below.
  2. Complete the ChatGPT sign-in and authorization prompts.
  3. Ask your assistant to use DeepWakeLabs and include the test settings and source links.

Connection name: DeepWakeLabs Research

This chatgpt.site address is the registered AI connection for deepwakelabs.com. Use the full address shown, even though it contains “preview.” Connection support depends on your AI client.

Streamable HTTP · Read-only data tools · v1.1.2

NO SIGN-IN NEEDED

Explore the published work.

The same evidence is available to read on the website.

Using an AI-enabled browser?

Checking browser support…

WebMCP lets a compatible browser agent discover the tools on this page. Data calls still require sign-in. On the benchmark page, two additional tools can read or adjust the visible filters without changing saved results.

Questions to start with

Advanced: try a data tool Requires an authenticated Site session

Inspect the actual response.

This tester uses your browser’s Site session. Connecting an AI client does not necessarily sign this browser in. If access is unavailable, you can still use the public links above.

No request is sent until you run a tool.

Choose a tool and run it to inspect its response.

9 DATA TOOLS · TWO REFERENCE RESOURCES

Built around the evidence.

Results include stable IDs, source links, capture precision, and recorded settings. Null values stay unknown. Drafts, private research files, and raw benchmark answers are excluded.

Explore the labget_lab_overview

Discover benchmark campaigns, released articles, revision, and evidence limitations. Start here to obtain exact campaign IDs. Catalog arrays are paginated independently; follow their nextOffset values. This is a curated snapshot, not a live benchmark runner.

{
  "type": "object",
  "properties": {
    "campaign_offset": {
      "type": "integer",
      "description": "Campaign catalog offset; default 0.",
      "minimum": 0,
      "maximum": 1000000
    },
    "article_offset": {
      "type": "integer",
      "description": "Article catalog offset; default 0.",
      "minimum": 0,
      "maximum": 1000000
    },
    "limit": {
      "type": "integer",
      "description": "Maximum entries per catalog; default 20.",
      "minimum": 1,
      "maximum": 25
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [],
  "additionalProperties": false
}
Search benchmark resultssearch_benchmarks

Search recorded results by model or quant and optionally campaign. Returns paginated measurements with runtime, capture precision, source links, missing values, and completeness. Sorting is descriptive and is not a cross-campaign ranking.

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "description": "Case-insensitive words in a model, publisher, or quant.",
      "minLength": 1,
      "maxLength": 160
    },
    "campaign_id": {
      "type": "string",
      "description": "Exact campaign ID from get_lab_overview.",
      "minLength": 1,
      "maxLength": 160
    },
    "status": {
      "type": "string",
      "description": "Exact evidence status, such as Complete or Historical.",
      "minLength": 1,
      "maxLength": 160
    },
    "sort": {
      "type": "string",
      "enum": [
        "source",
        "score",
        "speed",
        "memory",
        "name"
      ]
    },
    "offset": {
      "type": "integer",
      "description": "Zero-based offset; default 0.",
      "minimum": 0,
      "maximum": 1000000
    },
    "limit": {
      "type": "integer",
      "description": "Page size; default 10.",
      "minimum": 1,
      "maximum": 25
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [],
  "additionalProperties": false
}
Read exact resultsget_benchmark_results

Retrieve 1–10 exact result IDs in one call, including recorded configurations, capture timestamps, software versions, evidence status, and source citations. Missing IDs fail explicitly.

{
  "type": "object",
  "properties": {
    "result_ids": {
      "type": "array",
      "items": {
        "type": "string",
        "description": "Exact result ID.",
        "minLength": 1,
        "maxLength": 160
      },
      "minItems": 1,
      "maxItems": 10,
      "uniqueItems": true
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [
    "result_ids"
  ],
  "additionalProperties": false
}
Compare recorded configurationscompare_benchmarks

Compare 2–6 result IDs side by side. Reports differences in campaign, workload, thinking, cache, runtime, memory basis, and completeness. Cross-campaign data is descriptive only; no universal winner or predicted performance is calculated.

{
  "type": "object",
  "properties": {
    "result_ids": {
      "type": "array",
      "items": {
        "type": "string",
        "description": "Exact result ID.",
        "minLength": 1,
        "maxLength": 160
      },
      "minItems": 2,
      "maxItems": 6,
      "uniqueItems": true
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [
    "result_ids"
  ],
  "additionalProperties": false
}
Read a test profileget_campaign

Read a campaign’s benchmark edition, workload, configuration, capture precision, and source. Result IDs are paginated; follow resultPagination.nextOffset. Use search_benchmarks for full records.

{
  "type": "object",
  "properties": {
    "campaign_id": {
      "type": "string",
      "description": "Exact campaign ID.",
      "minLength": 1,
      "maxLength": 160
    },
    "offset": {
      "type": "integer",
      "description": "Result ID offset; default 0.",
      "minimum": 0,
      "maximum": 1000000
    },
    "limit": {
      "type": "integer",
      "description": "Result IDs per page; default 25.",
      "minimum": 1,
      "maximum": 100
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [
    "campaign_id"
  ],
  "additionalProperties": false
}
Find published analysissearch_articles

Search only the articles released on this Site. Returns source URLs, excerpts, dates, and linked benchmark campaigns. Article and video dates are not test capture dates.

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "description": "Words to find in article titles or content.",
      "minLength": 1,
      "maxLength": 160
    },
    "offset": {
      "type": "integer",
      "description": "Zero-based offset; default 0.",
      "minimum": 0,
      "maximum": 1000000
    },
    "limit": {
      "type": "integer",
      "description": "Page size; default 5.",
      "minimum": 1,
      "maximum": 10
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [],
  "additionalProperties": false
}
Read a published articleget_article

Read a released article by exact slug. Content is paginated without omission: follow contentPagination.nextOffset until null before treating the article as complete. Retains text, tables, sources, and test metadata. Drafts are excluded; claims reflect labeled dates.

{
  "type": "object",
  "properties": {
    "slug": {
      "type": "string",
      "description": "Article slug from search_articles or get_lab_overview.",
      "minLength": 1,
      "maxLength": 120
    },
    "offset": {
      "type": "integer",
      "description": "Content offset in Unicode characters; use returned nextOffset. Default 0.",
      "minimum": 0,
      "maximum": 10000000
    },
    "limit": {
      "type": "integer",
      "description": "Content characters per response; default 12000.",
      "minimum": 1,
      "maximum": 20000
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [
    "slug"
  ],
  "additionalProperties": false
}
Inspect GPU memory estimatesget_memory_screening

Read projected 12/16 GiB memory screenings, reserve assumptions, and paginated candidate runtimes. These measurements came from an RTX 5090; physical 12/16 GB qualification and quality rankings are not implied.

{
  "type": "object",
  "properties": {
    "budget_gib": {
      "type": "integer",
      "enum": [
        12,
        16
      ]
    },
    "offset": {
      "type": "integer",
      "description": "Capture metadata offset; default 0.",
      "minimum": 0,
      "maximum": 1000000
    },
    "limit": {
      "type": "integer",
      "description": "Capture metadata entries per page; default 10.",
      "minimum": 1,
      "maximum": 25
    },
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [],
  "additionalProperties": false
}
Understand the evidenceget_methodology

Read the rules for interpreting quality scores, generation speed, memory units, missing data, capture dates, and regrades. Use before comparing different benchmark profiles.

{
  "type": "object",
  "properties": {
    "expected_revision": {
      "type": "string",
      "description": "Optional meta.datasetRevision from a previous response. Use when paging or comparing across calls; a changed publication returns STALE_SNAPSHOT instead of mixing revisions.",
      "minLength": 1,
      "maxLength": 64
    }
  },
  "required": [],
  "additionalProperties": false
}

DEEPWAKELABS