Methodology
How VyDex decides what enters the ledger, how claims are judged, and what each public label means.
- Current Version
- v1.0.0
- Effective From
- Version Type
- Major
- Entry linkage
- Entry pages link to the Methodology Version used for that entry record.
Inclusion Rule
VyDex includes claims that appear to cross, challenge, weaken, verify, or materially change a frontier threshold.
Ordinary product launches, funding rounds, routine trials, cool demos, and general news are excluded unless they change the frontier evidence picture.
Inclusion Standard
A claim belongs in VyDex when it meaningfully changes the frontier evidence picture.
Does the claim cross, challenge, weaken, verify, or materially change a frontier threshold?
Is there enough source material to explain what is being claimed?
Can VyDex state what the evidence supports without repeating a headline uncritically?
Would tracking this help users understand how a frontier area is changing over time?
- Included Example
A new evaluation shows agents reaching a clearly longer software-task horizon, with enough methodology detail to judge the scope.
- Excluded Example
A company announces a new model, product, or funding round without evidence that a frontier threshold changed.
Claim Appraisal
Every serious entry asks what is claimed, what evidence exists, what the evidence measures, what conditions shaped it, what uncertainty remains, and what conclusion is warranted.
What exactly is being claimed?
What source material supports it?
What does the evidence actually measure?
What setup, baseline, comparison, or constraint changes interpretation?
What uncertainty remains?
What can VyDex responsibly conclude?
Public Labels
These labels appear on entry cards and Entry Pages. They help users interpret the record; they are not popularity signals or importance scores.
Claim Status
| Status | Meaning | UI Treatment |
|---|---|---|
| Confirmed | The stated claim is strongly supported within its stated scope by durable evidence, such as replication, external verification, official records, strong artifacts, or multiple converging sources. | Neutral label |
| Supported | The claim is reasonably supported by current evidence, but not strong enough to call confirmed. | Neutral label |
| Provisional | The claim is plausible and worth documenting, but early, limited, narrow, source-dependent, or not yet well verified. | Neutral label |
| Reported But Unverified | The claim has been reported by a source worth tracking, but VyDex does not yet have enough primary evidence, technical detail, or independent confirmation to treat it as supported. | Quiet caution |
| Disputed | Credible disagreement, conflicting evidence, failed replication attempts, methodology concerns, or strong caveats materially challenge the claim. | Medium caution |
| Failed / Retracted | The claim has failed, been retracted, been corrected in a way that breaks the original claim, or no longer supports the frontier interpretation. | Strongest negative treatment |
Evidence Strength
Evidence Strength describes how strong the current support is for the entry’s stated claim within its stated scope. It is not probability, confidence, importance, or ranking.
| Level | Score | Meaning | Typical Evidence |
|---|---|---|---|
| Thin | 1 | Limited support, unclear evidence, narrow detail, or heavy dependence on one source. | Early developer/vendor claim, unclear demo, weak artifact, rumor/leak worth tracking, or media report without strong technical detail. |
| Moderate | 2 | Enough evidence to document responsibly, but still with meaningful limits. | Primary source with technical detail, clear preprint, credible independent coverage, public demo with artifacts, or official claim that supports the basic event. |
| Strong | 3 | Solid evidence beyond a bare assertion. | Primary evidence plus independent analysis, public logs, released artifacts, credible replication attempts, official filings, deployment evidence, or detailed benchmark methodology. |
| Very Strong | 4 | Unusually strong support and low ambiguity within the stated scope. | Independent replication, peer review plus strong artifacts, external audit, formal verification, legal/government confirmation, durable public records, or multiple independent confirmations. |
Review Status
Review Status describes whether the entry needs active follow-up.
| Status | Meaning | When Used |
|---|---|---|
| Stable | No active follow-up is currently scheduled. | Durable evidence, settled scope, or no known fast-moving uncertainty. |
| Follow-Up Needed | The entry should be checked again because the evidence may change or the claim is still volatile. | Disputed claims, preprints, vendor/developer claims, reported-but-unverified claims, missing replication, recently changed evidence, unclear artifacts, or scheduled review triggers. |
- Review Reason
Review Reason explains why an Entry is marked Follow-Up Needed and identifies the evidence, uncertainty, or review trigger that warrants another review.
Entry State
| State | Meaning |
|---|---|
| Main Entry | The entry is part of the current public evidence ledger. |
| Removed | The entry no longer meets criteria or no longer supports the frontier interpretation, but may remain explainable through Changelog or later version history behavior. |
Entry Fields
Frontier Delta
Frontier Delta explains what changed from the previous relevant frontier state.
- Previous Frontier
The previous relevant frontier state.
- New Claim / Result
What the new claim or result says.
- Delta
The meaningful change from the previous frontier to the new claim or result.
Significance
- Confirmed Significance
What the current evidence actually supports within the entry’s stated scope. Required for every main entry.
- Potential Significance If Confirmed
Why the claim would matter if later evidence strengthens, verifies, replicates, audits, or confirms it. Used for early, disputed, preprint-only, vendor/developer-led, reported-but-unverified, low-evidence, or easy-to-overread claims.
Caveats
Caveats explain what could make an entry easy to overread.
Narrow scope.
Weak generalization.
Unclear baseline.
Special setup.
Missing replication.
Vendor-controlled evidence.
Disputed methodology.
Limited deployment evidence.
Source uncertainty.
Sources and Evidence Types
Sources must show what they are and what VyDex used them for. Raw unexplained links are not enough.
Evidence Types
| Evidence Type | Meaning |
|---|---|
| Preprint | A research paper or working paper made public before completing formal peer review. |
| Peer-Reviewed Paper | A research paper published by a journal or conference after formal peer review. |
| Independent Replication | A study or test by a separate party that attempts to reproduce or verify a reported result. |
| Official Claim | A formal claim made by a government body, regulator, court, public institution, or other authority acting in an official capacity. |
| Developer / Vendor Claim | A claim made by the people or organization that developed, operates, or sells the system, product, or service being discussed. |
| Benchmark Result | A reported score or outcome from a defined benchmark, test suite, or evaluation procedure. |
| Technical Artifact | Code, datasets, model weights, logs, executables, or other technical material that can be directly inspected. |
| Government Report | A report published by a government body or agency. |
| Court Filing | A document filed in a court or other formal legal proceeding. |
| Audit | A documented audit or formal assessment of a system, process, claim, or set of controls. |
| Media Report | A report published by a journalistic or news organization. |
| Leaked / Internal Claim | A claim or document originating from non-public internal material or an unauthorized disclosure. |
Used For
- Definition
Used For explains what information VyDex took from a source.
- Public Statement
Each source says what VyDex used it for.
- Example
Used For: Benchmark setup, measured task horizon, and stated limitations.
Source Roles
| Source Role | Meaning |
|---|---|
| Primary Evidence | The main direct evidence for the Entry’s stated claim. |
| Independent Replication | Evidence from a separate party attempting to reproduce or verify the result. |
| Official Record | A durable official document, filing, ruling, government publication, or equivalent record. |
| Strong Artifact | A substantive artifact such as logs, code, benchmark records, datasets, recordings, or technical documentation. |
| Context Source | A source used primarily for background, comparison, scope, or interpretation. |
| Media Report | Journalism or reporting used to establish or contextualize facts when stronger direct evidence is unavailable or incomplete. |
Source Role and Evidence Type must not be confused. A government report is an Evidence Type, while Official Record may be its role in a particular Entry. A technical artifact may be Primary Evidence in one Entry and supporting context in another.
Source Ordering
Sources are ordered by evidence role. Primary evidence, replication, official records, and strong artifacts appear before weaker context sources or media reports.
Dates and Evidence Monitoring
Date Fields
| Field | Meaning |
|---|---|
| Date Happened | When the event, result, deployment, incident, benchmark run, law, publication, or frontier change actually happened, if known. |
| Date Disclosed | When it became public, was announced, reported, published, filed, or otherwise disclosed. |
| Date Added | When VyDex added the entry. |
| Date Updated | When VyDex last materially changed the entry. |
| Date Last Checked | When VyDex last checked the entry’s evidence state. |
| Next Check Date | When VyDex plans to check the entry again, if follow-up is needed or scheduled. |
Evidence Monitoring
VyDex uses evidence monitoring so unstable or fast-moving claims are not forgotten, while stable entries are not rechecked without a reason.
New evidence.
Independent replication.
Peer review.
Failed replication.
Source changes.
User correction reports.
Related Topic Trail movement.
Scheduled light audit.
Topic Trails and Domains
Topic Trails
A Topic Trail is a controlled ongoing frontier storyline that entries belong to.
Each entry has a Topic Trail.
Most entries have one primary Topic Trail.
Secondary Topic Trails are added only when they materially improve interpretation.
- Naming Rule
Topic Trail names should be specific enough to understand, broad enough to hold multiple entries, neutral enough not to assume the outcome, stable enough to last, and searchable enough for normal users.
- Good Examples
AI agents in software engineering.
AI in research mathematics.
AI cyber capability.
Inference cost and efficiency collapse.
Medical AI moving toward clinical authority.
AI systems in scientific discovery.
AI governance and state control.
- Bad Examples
AI is taking over coding.
GPT-5.
Benchmarks.
Robotics.
FireRed completion, unless it becomes a repeated evaluation storyline.
Domains
| Domain | Scope |
|---|---|
| AI Capabilities | Models, agents, reasoning, coding, multimodal systems, autonomy, AI tool use, and general AI capability claims. |
| AI Evaluation | Benchmarks, eval failures, benchmark saturation, new evaluation methods, measurement problems, audit results, and claims about what current systems can or cannot do. |
| Robotics | Humanoid robots, manipulation, locomotion, drones, embodied agents, physical-world autonomy, factories, and field deployment. |
| Biology | Biology research, drug discovery, medical AI, clinical deployment, diagnostics, trials, wet-lab automation, and health-related frontier claims. |
| Mathematics | Mathematical proof, formal verification, theorem proving, logic, symbolic reasoning, and research-level math progress. |
| Physical Sciences | Physics, chemistry, materials science, superconductors, batteries, energy materials, nanotech, and non-biomedical lab science. |
| Cybersecurity | Vulnerability discovery, exploit generation, cyber ranges, autonomous cyber agents, defensive systems, offensive capability claims, and cyber-risk evaluations. |
| Hardware | Chips, GPUs, accelerators, datacenters, training compute, inference infrastructure, memory systems, networking, and hardware efficiency. |
| Economics | Inference cost, training cost, cost collapse, labor productivity, automation economics, business deployment, and economic-scale effects. |
| Governance | Regulation, export controls, court rulings, safety institutes, state review, standards, liability, legal authority, and public-sector policy. |
| National Security | Military AI, autonomous weapons, intelligence use, battlefield deployment, national-security claims, defense procurement, and strategic-security implications. |
| Space | Space systems, launch, satellites, and large physical frontier space-related projects. |
Entry Titles
- Rule
Entry titles must be specific, evidence-bounded, and non-clickbait.
- Pattern
actor/system + crossed/failed/approached/similar verb + concrete threshold + scope/constraint
- Hype-Word Rule
Titles avoid vague hype words such as revolutionary, insane, game-changing, stunning, massive, breakthrough, and unprecedented unless the entry immediately defines and justifies the term.
- Examples
Fable 5 completes Pokémon FireRed using screenshots only
METR finds frontier agents reaching ~50-minute software-task horizons
Versioning
VyDex uses versioned methodology so users can see which rules applied to an entry record.
| Change Type | When It Changes |
|---|---|
| Major | A methodology change could alter how entries are judged, such as changing inclusion standards, redefining statuses, changing Evidence Strength, changing ranking logic meaningfully, or changing the main entry vs watchlist boundary. |
| Minor | A new category, label, clarification, source role, evidence example, or review rule improves the methodology without overturning previous judgments. |
| Patch | Wording fixes, typo corrections, clarified examples, repaired links, or changes that do not materially affect judgment. |
Current Stage 1 methodology version: v1.0.0