Methodology

The measurement contract.

Fulcrum publishes the rules it measures by, so a number can be checked rather than believed. This is the contract the software enforces, in the order the software enforces it.

Scope

A longitudinal record of one operator, not a benchmark.

A Fulcrum dataset describes one team's operating conditions and results. It is not a population benchmark, a staffing model, or a claim about an average engineer. The record is most useful when it preserves the operator conditions, the estimate provenance, and the capability boundaries that produced the result, which is why every field below carries its class.

Evidence classes

Every public field belongs to exactly one class.

Measured execution

Captured dates, elapsed agent-session minutes, token counts with exact provenance, provider and model identifiers, and operator time recorded during the task.

Estimated counterfactual

How long equivalent work would take without the agentic implementation path. Estimates are ranges with a named basis and author.

Operator conditions

Experience, product authority, work cadence, tooling, subscriptions, and parallel orchestration. Boundary conditions, not multiplier coefficients.

Capability expansion

Work that required substantial new learning or was not reasonably feasible unaided. Classified, never assigned invented or infinite hour estimates.

Canonical identity

One task, one record, forever.

  • New records use a client-generated client_record_id that is unique within a user account.
  • Repeating a create or completion request with the same id returns or updates the same canonical record.
  • Existing records keep their server UUID as canonical identity and are tagged with legacy estimate provenance.
  • A date, task, and metrics fingerprint may identify reconciliation candidates, but it is never a hard identity rule: two legitimate tasks can share those values.
  • Corrections preserve the original row and use an explicit supersession relationship or correction ledger. Public outputs contain only the effective record.
Counterfactual estimates

Reproducible historically, richer going forward.

human_estimate_hours remains the qualified-senior-equivalent estimate so existing records and formulas stay reproducible. New capture can also record an operator-unassisted range: personal_estimate_low_hours, personal_estimate_likely_hours, and personal_estimate_high_hours. The range must be monotonic, and it stays null when the work is classified as not reasonably feasible unaided.

Provenance recorded with every estimate

  • Basis: legacy_unspecified or qualified_senior
  • Author: agent, operator, or joint
  • Operator feasibility: can_do, can_do_with_learning, or not_reasonably_feasible
  • Capability class: time_compression, skill_extension, or otherwise_infeasible
  • A concise rationale whenever a personal range or a capability-expansion classification is supplied

Historical estimates are never multiplied in place. A reviewed sample or a sensitivity analysis is kept as a separately versioned interpretation.

Time and token provenance

Minutes that mean one thing each.

  • agent_session_minutes is the elapsed time for one agent task. The legacy field claude_minutes remains its compatibility name until a versioned API removes it.
  • Period totals sum agent sessions and can exceed 24 hours per calendar day when sessions overlap.
  • wall_clock_span_minutes is the elapsed start-to-finish span for a task. It is not summed as a claim of unique calendar time unless overlapping intervals have been unioned.
  • Operator time is the sum of prompt composition, active supervision, correction, and review minutes.
  • Tokens are exact, estimated, or unavailable. Estimated tokens include a named method or rate; unavailable values are null or zero with unavailable provenance, never described as measured.
Published metrics

Let H be qualified-senior hours, A agent-session minutes, O operator minutes.

Aggregate execution factor

60 × sum(H) / sum(A)

The primary execution metric. A ratio of sums; it is not the arithmetic mean of task factors.

Median task factor

median(60 × H / A)

Describes the middle individual task.

Arithmetic task mean

mean(60 × H / A)

Retained and labeled as the task mean. Never labeled simply “average leverage” when the aggregate ratio is also present.

Aggregate operator leverage

60 × sum(H) / sum(O)

Calculated only over records with complete operator-time coverage. Every result includes the covered record count and percentage.

Calendar compression

60 × H / wall_clock_span_minutes

At period level the denominator is the union of captured task intervals. Summing overlapping spans is not unique calendar time.

Capability expansion

counts and shares by feasibility and class

No numeric leverage factor is ever assigned on the basis of an invented unassisted duration.

Publication eligibility

A record is public only when all of these are true.

  • It is not deleted or superseded.
  • It is confirmed rather than draft.
  • Its realization is productive, or it is a legacy record without a contradictory realization.
  • Its publication review is approved.
  • It has a reviewed public-safe task restatement.
  • Its numeric source fields pass structural and formula validation.

Raw task text never crosses the publication boundary.

The snapshot contract

One immutable publication build emits leverage_records.csv, leverage_records.json, summary.json, timeseries.json, METHODOLOGY.json, CORRECTIONS.json, and MANIFEST.json, plus a ZIP with the same files and human-readable documentation. The manifest carries the methodology and schema versions, the snapshot id and canonical high-water timestamp, the record count and date range, each published metric with its formula identifier, the evidence and provenance coverage counts, and a SHA-256 for every sibling artifact. The website, the download, and the generated daily rollups consume the same snapshot, and a build fails when counts, totals, formulas, or hashes do not reconcile.

Measure it before you argue about it.

One annual license installs Fulcrum in your environment: the API, the dashboard, the CLI, and the MCP server your agents already know how to call. $1,199 per year. Included for a year with every Vantalect engagement.

Already licensed? Sign in to Fulcrum.