Skip to main content

Methodology

How VyDex decides what enters the ledger, how claims are judged, and what each public label means.

Current Version
v1.0.0
Effective From
Version Type
Major
Entry linkage
Entry pages link to the Methodology Version used for that entry record.

Inclusion Rule

VyDex includes claims that appear to cross, challenge, weaken, verify, or materially change a frontier threshold.

Ordinary product launches, funding rounds, routine trials, cool demos, and general news are excluded unless they change the frontier evidence picture.

Inclusion Standard

A claim belongs in VyDex when it meaningfully changes the frontier evidence picture.

  • Does the claim cross, challenge, weaken, verify, or materially change a frontier threshold?

  • Is there enough source material to explain what is being claimed?

  • Can VyDex state what the evidence supports without repeating a headline uncritically?

  • Would tracking this help users understand how a frontier area is changing over time?

Included Example

A new evaluation shows agents reaching a clearly longer software-task horizon, with enough methodology detail to judge the scope.

Excluded Example

A company announces a new model, product, or funding round without evidence that a frontier threshold changed.

Claim Appraisal

Every serious entry asks what is claimed, what evidence exists, what the evidence measures, what conditions shaped it, what uncertainty remains, and what conclusion is warranted.

  1. What exactly is being claimed?

  2. What source material supports it?

  3. What does the evidence actually measure?

  4. What setup, baseline, comparison, or constraint changes interpretation?

  5. What uncertainty remains?

  6. What can VyDex responsibly conclude?

Public Labels

These labels appear on entry cards and Entry Pages. They help users interpret the record; they are not popularity signals or importance scores.

Claim Status

StatusMeaningUI Treatment
Confirmed

The stated claim is strongly supported within its stated scope by durable evidence, such as replication, external verification, official records, strong artifacts, or multiple converging sources.

Neutral label
Supported

The claim is reasonably supported by current evidence, but not strong enough to call confirmed.

Neutral label
Provisional

The claim is plausible and worth documenting, but early, limited, narrow, source-dependent, or not yet well verified.

Neutral label
Reported But Unverified

The claim has been reported by a source worth tracking, but VyDex does not yet have enough primary evidence, technical detail, or independent confirmation to treat it as supported.

Quiet caution
Disputed

Credible disagreement, conflicting evidence, failed replication attempts, methodology concerns, or strong caveats materially challenge the claim.

Medium caution
Failed / Retracted

The claim has failed, been retracted, been corrected in a way that breaks the original claim, or no longer supports the frontier interpretation.

Strongest negative treatment

Evidence Strength

Evidence Strength describes how strong the current support is for the entry’s stated claim within its stated scope. It is not probability, confidence, importance, or ranking.

LevelScoreMeaningTypical Evidence
Thin1

Limited support, unclear evidence, narrow detail, or heavy dependence on one source.

Early developer/vendor claim, unclear demo, weak artifact, rumor/leak worth tracking, or media report without strong technical detail.

Moderate2

Enough evidence to document responsibly, but still with meaningful limits.

Primary source with technical detail, clear preprint, credible independent coverage, public demo with artifacts, or official claim that supports the basic event.

Strong3

Solid evidence beyond a bare assertion.

Primary evidence plus independent analysis, public logs, released artifacts, credible replication attempts, official filings, deployment evidence, or detailed benchmark methodology.

Very Strong4

Unusually strong support and low ambiguity within the stated scope.

Independent replication, peer review plus strong artifacts, external audit, formal verification, legal/government confirmation, durable public records, or multiple independent confirmations.

Review Status

Review Status describes whether the entry needs active follow-up.

StatusMeaningWhen Used
Stable

No active follow-up is currently scheduled.

Durable evidence, settled scope, or no known fast-moving uncertainty.

Follow-Up Needed

The entry should be checked again because the evidence may change or the claim is still volatile.

Disputed claims, preprints, vendor/developer claims, reported-but-unverified claims, missing replication, recently changed evidence, unclear artifacts, or scheduled review triggers.

Review Reason

Review Reason explains why an Entry is marked Follow-Up Needed and identifies the evidence, uncertainty, or review trigger that warrants another review.

Entry State

StateMeaning
Main Entry

The entry is part of the current public evidence ledger.

Removed

The entry no longer meets criteria or no longer supports the frontier interpretation, but may remain explainable through Changelog or later version history behavior.

Entry Fields

Frontier Delta

Frontier Delta explains what changed from the previous relevant frontier state.

Previous Frontier

The previous relevant frontier state.

New Claim / Result

What the new claim or result says.

Delta

The meaningful change from the previous frontier to the new claim or result.

Significance

Confirmed Significance

What the current evidence actually supports within the entry’s stated scope. Required for every main entry.

Potential Significance If Confirmed

Why the claim would matter if later evidence strengthens, verifies, replicates, audits, or confirms it. Used for early, disputed, preprint-only, vendor/developer-led, reported-but-unverified, low-evidence, or easy-to-overread claims.

Caveats

Caveats explain what could make an entry easy to overread.

  • Narrow scope.

  • Weak generalization.

  • Unclear baseline.

  • Special setup.

  • Missing replication.

  • Vendor-controlled evidence.

  • Disputed methodology.

  • Limited deployment evidence.

  • Source uncertainty.

Sources and Evidence Types

Sources must show what they are and what VyDex used them for. Raw unexplained links are not enough.

Evidence Types

Evidence TypeMeaning
Preprint

A research paper or working paper made public before completing formal peer review.

Peer-Reviewed Paper

A research paper published by a journal or conference after formal peer review.

Independent Replication

A study or test by a separate party that attempts to reproduce or verify a reported result.

Official Claim

A formal claim made by a government body, regulator, court, public institution, or other authority acting in an official capacity.

Developer / Vendor Claim

A claim made by the people or organization that developed, operates, or sells the system, product, or service being discussed.

Benchmark Result

A reported score or outcome from a defined benchmark, test suite, or evaluation procedure.

Technical Artifact

Code, datasets, model weights, logs, executables, or other technical material that can be directly inspected.

Government Report

A report published by a government body or agency.

Court Filing

A document filed in a court or other formal legal proceeding.

Audit

A documented audit or formal assessment of a system, process, claim, or set of controls.

Media Report

A report published by a journalistic or news organization.

Leaked / Internal Claim

A claim or document originating from non-public internal material or an unauthorized disclosure.

Used For

Definition

Used For explains what information VyDex took from a source.

Public Statement

Each source says what VyDex used it for.

Example

Used For: Benchmark setup, measured task horizon, and stated limitations.

Source Roles

Source RoleMeaning
Primary Evidence

The main direct evidence for the Entry’s stated claim.

Independent Replication

Evidence from a separate party attempting to reproduce or verify the result.

Official Record

A durable official document, filing, ruling, government publication, or equivalent record.

Strong Artifact

A substantive artifact such as logs, code, benchmark records, datasets, recordings, or technical documentation.

Context Source

A source used primarily for background, comparison, scope, or interpretation.

Media Report

Journalism or reporting used to establish or contextualize facts when stronger direct evidence is unavailable or incomplete.

Source Role and Evidence Type must not be confused. A government report is an Evidence Type, while Official Record may be its role in a particular Entry. A technical artifact may be Primary Evidence in one Entry and supporting context in another.

Source Ordering

Sources are ordered by evidence role. Primary evidence, replication, official records, and strong artifacts appear before weaker context sources or media reports.

Dates and Evidence Monitoring

Date Fields

FieldMeaning
Date Happened

When the event, result, deployment, incident, benchmark run, law, publication, or frontier change actually happened, if known.

Date Disclosed

When it became public, was announced, reported, published, filed, or otherwise disclosed.

Date Added

When VyDex added the entry.

Date Updated

When VyDex last materially changed the entry.

Date Last Checked

When VyDex last checked the entry’s evidence state.

Next Check Date

When VyDex plans to check the entry again, if follow-up is needed or scheduled.

Evidence Monitoring

VyDex uses evidence monitoring so unstable or fast-moving claims are not forgotten, while stable entries are not rechecked without a reason.

  • New evidence.

  • Independent replication.

  • Peer review.

  • Failed replication.

  • Source changes.

  • User correction reports.

  • Related Topic Trail movement.

  • Scheduled light audit.

Topic Trails and Domains

Topic Trails

A Topic Trail is a controlled ongoing frontier storyline that entries belong to.

  • Each entry has a Topic Trail.

  • Most entries have one primary Topic Trail.

  • Secondary Topic Trails are added only when they materially improve interpretation.

Naming Rule

Topic Trail names should be specific enough to understand, broad enough to hold multiple entries, neutral enough not to assume the outcome, stable enough to last, and searchable enough for normal users.

Good Examples
  • AI agents in software engineering.

  • AI in research mathematics.

  • AI cyber capability.

  • Inference cost and efficiency collapse.

  • Medical AI moving toward clinical authority.

  • AI systems in scientific discovery.

  • AI governance and state control.

Bad Examples
  • AI is taking over coding.

  • GPT-5.

  • Benchmarks.

  • Robotics.

  • FireRed completion, unless it becomes a repeated evaluation storyline.

Domains

DomainScope
AI Capabilities

Models, agents, reasoning, coding, multimodal systems, autonomy, AI tool use, and general AI capability claims.

AI Evaluation

Benchmarks, eval failures, benchmark saturation, new evaluation methods, measurement problems, audit results, and claims about what current systems can or cannot do.

Robotics

Humanoid robots, manipulation, locomotion, drones, embodied agents, physical-world autonomy, factories, and field deployment.

Biology

Biology research, drug discovery, medical AI, clinical deployment, diagnostics, trials, wet-lab automation, and health-related frontier claims.

Mathematics

Mathematical proof, formal verification, theorem proving, logic, symbolic reasoning, and research-level math progress.

Physical Sciences

Physics, chemistry, materials science, superconductors, batteries, energy materials, nanotech, and non-biomedical lab science.

Cybersecurity

Vulnerability discovery, exploit generation, cyber ranges, autonomous cyber agents, defensive systems, offensive capability claims, and cyber-risk evaluations.

Hardware

Chips, GPUs, accelerators, datacenters, training compute, inference infrastructure, memory systems, networking, and hardware efficiency.

Economics

Inference cost, training cost, cost collapse, labor productivity, automation economics, business deployment, and economic-scale effects.

Governance

Regulation, export controls, court rulings, safety institutes, state review, standards, liability, legal authority, and public-sector policy.

National Security

Military AI, autonomous weapons, intelligence use, battlefield deployment, national-security claims, defense procurement, and strategic-security implications.

Space

Space systems, launch, satellites, and large physical frontier space-related projects.

Entry Titles

Rule

Entry titles must be specific, evidence-bounded, and non-clickbait.

Pattern

actor/system + crossed/failed/approached/similar verb + concrete threshold + scope/constraint

Hype-Word Rule

Titles avoid vague hype words such as revolutionary, insane, game-changing, stunning, massive, breakthrough, and unprecedented unless the entry immediately defines and justifies the term.

Examples
  • Fable 5 completes Pokémon FireRed using screenshots only

  • METR finds frontier agents reaching ~50-minute software-task horizons

Versioning

VyDex uses versioned methodology so users can see which rules applied to an entry record.

Change TypeWhen It Changes
Major

A methodology change could alter how entries are judged, such as changing inclusion standards, redefining statuses, changing Evidence Strength, changing ranking logic meaningfully, or changing the main entry vs watchlist boundary.

Minor

A new category, label, clarification, source role, evidence example, or review rule improves the methodology without overturning previous judgments.

Patch

Wording fixes, typo corrections, clarified examples, repaired links, or changes that do not materially affect judgment.

Current Stage 1 methodology version: v1.0.0