ArcProof · Skill Foundations · Certifications

ArcProof Research Institute · September 3, 2026

The ArcProof Verification Framework™

A four-stage methodology for evidence-based, predictively valid skill verification — from construct definition to longitudinal outcome validation.

Version 1.0 (working draft)

Author of record: ArcProof Research Institute (a research initiative of SkillUpArc LLC)

Published: September 3, 2026

© 2026 SkillUpArc LLC. The ArcProof Verification Framework™ is a trademark of SkillUpArc LLC.

Abstract

As artificial intelligence makes instructional content abundant and lowers the cost of fabricating skill claims, the credibility of self-reported and completion-based credentials erodes. This paper introduces the ArcProof Verification Framework — a methodology for producing skill credentials whose meaning is grounded in evidence, protected by scoring integrity, and validated against real-world outcomes. The framework synthesizes established work in evidence-centered assessment design (Mislevy, Steinberg, & Almond, 2003) and argument-based validity theory (Kane, 2013; Messick, 1989) into a four-stage model suited to an AI-mediated labor market. Its distinguishing contribution is a longitudinal validation stage that treats a credential's predictive relationship to employment outcomes as an empirical claim to be tested over time rather than asserted at issuance.

1. The problem: verification in an era of abundant content and cheap claims

For most of the history of credentialing, a certificate of completion carried a workable signal: completing a course of study was costly in time and effort, so completion loosely indicated capability. Two developments have weakened that signal. First, AI-mediated instruction is driving the marginal cost of high-quality content toward zero, so completion becomes both more common and less differentiating. Second, the same technologies make it inexpensive to generate a convincing resume, satisfy a recognition-based assessment, or fabricate a credential — so self-reported capability becomes difficult to distinguish from demonstrated capability.

The consequence is a widening gap between what a credential asserts ("this learning occurred") and what a decision-maker needs to know ("this person can perform this work"). Closing that gap requires returning to first principles of educational measurement: a credential is only as sound as the evidentiary argument that connects a person's performance to a defensible claim about their capability (Kane, 2013; Mislevy & Haertel, 2006).

2. Theoretical foundations

The framework rests on two well-established bodies of work, neither proprietary.

Evidence-centered design (ECD). ECD reframes assessment as the construction of an evidentiary argument: from claims about what a person knows or can do, to the evidence that would warrant those claims, to the tasks capable of eliciting that evidence (Mislevy, Steinberg, & Almond, 2003; Mislevy, Almond, & Lukas, 2003). ECD's insistence that assessment begin with an explicit argument — rather than with items — is the design discipline underlying this framework.

Argument-based validity. Contemporary validity theory holds that it is not a test that is valid, but a specific interpretation and use of its scores, and that validity is established by evaluating the argument connecting scores to those interpretations and uses (Kane, 2006, 2013; Messick, 1989). The Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) codify the categories of evidence — content, response process, internal structure, relations to other variables, and consequences — that such an argument must marshal. For performance-based credentials in particular, the interplay of evidence and consequences is central (Messick, 1994).

The originality of the ArcProof Verification Framework lies not in these foundations but in their synthesis into a verification pipeline for AI-mediated skill claims, and specifically in (a) an integrity-assured scoring stage designed to resist gaming, and (b) a longitudinal outcome-validation stage that makes predictive validity an ongoing empirical commitment.

3. The framework: four stages

The framework traces the progression — the arc — a claim travels to become verified, predictively meaningful capability. Each stage answers a distinct question in the evidentiary argument, and each depends on the one before it.

Stage 1 — Construct definition

What capability is being claimed, and does it matter? No score is issued without a defined construct mapped to real occupational demand. This anchors every credential to a specific, role-relevant claim rather than a generic label, and it is the point at which content-relevance evidence (AERA, APA, & NCME, 2014) is established.

Stage 2 — Evidence elicitation

What performance would warrant the claim? Capability is demonstrated through authentic, performance-based tasks that generate a durable evidence trail, consistent with ECD's task-model logic (Mislevy et al., 2003) and the performance-assessment tradition (Messick, 1994). Recognition-based items alone are treated as insufficient warrant.

Stage 3 — Integrity-assured scoring

Is the evidence evaluated consistently, and is the result resistant to manipulation? Evidence is scored against a fixed, criterion-referenced rubric, applied consistently and auditably, with a tamper-evident record. The design goal is that a score cannot be raised by consuming more content or by gaming the interface — only by demonstrating the construct. The internal mechanics of this stage are not disclosed here.

Stage 4 — Predictive validation

Does the credential predict what it claims to predict? Issued credentials are tracked longitudinally against real outcomes — hiring, advancement, earnings, retention — so that the relationship between a verified capability profile and downstream performance is established empirically and updated over time. This stage operationalizes "relations to other variables" evidence (AERA, APA, & NCME, 2014) as a continuing commitment rather than a one-time study. It is the stage that distinguishes a credential that is verifiable ("this was genuinely earned") from one that is predictive ("this reliably forecasts performance").

4. Integration: a single validity argument

The four stages compose one validity argument in Kane's (2013) sense. Construct definition specifies the claim; evidence elicitation and integrity-assured scoring warrant the scoring and generalization inferences; predictive validation warrants the extrapolation and decision inferences that justify real-world use. Because each stage depends on the prior one, weakness at any stage bounds the strength of the whole — a well-defined construct with gameable scoring yields an untrustworthy credential, and rigorous scoring with no outcome validation yields a credential that is verifiable but not yet predictive.

5. Distinction from existing approaches

The framework is defined partly by what it rejects. Completion certificates warrant only that instruction was consumed. Self-reported and endorsement-based signals warrant nothing beyond the claimant's or endorser's assertion. Issuer-vouched digital badges certify that some authority attests to a credential, but typically do not themselves run the assessment or validate outcomes. The ArcProof Verification Framework's contribution is to combine independent, evidence-based assessment with tamper-evident scoring and a standing outcome-validation commitment, applied to the individual credential-holder.

6. Applications

The framework is intended to govern any context where a capability claim must be trusted by a third party who did not observe the learning: individual professionals seeking portable proof, employers screening candidates, and workforce and education programs required to document verified outcomes for funders and auditors. In each case, the framework provides a common standard for what "verified" must mean.

7. Limitations and future validation

Consistent with the framework's own principles, its central claims are empirical and not yet fully established. Predictive validity (Stage 4) is, by definition, a function of accumulated longitudinal data; early credentials carry a validation argument that is sound in design but thin in evidence, and that strengthens only as outcome data accrue. The consistency and fairness of automated scoring (Stage 3) across populations require ongoing monitoring of the kind the Standards (AERA, APA, & NCME, 2014) prescribe. This paper states the framework; establishing its validity is the work of the years that follow — which is itself the point.

References

Suggested citation

ArcProof Research Institute. (2026). The ArcProof Verification Framework™: Evidence-based, predictively valid skill verification (Version 1.0). SkillUpArc LLC. https://www.skilluparc.com/framework

© 2026 SkillUpArc LLC. The ArcProof Verification Framework™ is a trademark of SkillUpArc LLC.