Methodology

How assessment and verification work

This page explains what we measure, how scores are produced, where AI is used, where a human decides, and how any result can be challenged. It is written to be read by the people being assessed, not only by the people hiring.

What we measure

We publish a multidimensional scorecard rather than a single mysterious number. Each dimension is scored separately, carries its own confidence value, and links to the evidence behind it.

DimensionWhat it meansEvidence
Role CompetencyKnowledge and applied understanding of the target role.Knowledge items and scenario tasks.
PracticalQuality of a real work simulation.Work artifact scored against a rubric, plus reviewer check.
CommunicationWritten and voice client communication.Situational responses and intro video transcript.
EnglishFunctional workplace English.Written, listening and speaking tasks.
AI FluencyEffective, safe and verified AI-assisted workflow.Scenario and practical exercise.
Portfolio EvidenceStrength and verifiability of prior work.Source files, version history, client confirmation.
ReputationPerformance inside Taskopedia.Reviews, completion, repeat work, disputes.
ReliabilityDelivery behaviour, not personality inference.On-time milestones, responsiveness, attendance where contracted.

How a score is calculated

A dimension score is the weighted mean of the scored items within it. Confidence is calculated separately, as the product of how much of the blueprint the evidence actually covers, how reliable the assessment instrument is, and whether a human reviewed the result.

dimension_score = Σ(item_score × item_weight) / Σ(item_weight)
confidence      = evidence_coverage × assessment_reliability × review_factor
published_score = round(dimension_score)

The weights are held in versioned configuration that is reviewed and published, never inside a prompt. A result is always pinned to the assessment version it was taken under, so changing a rubric never silently rewrites scores that already exist.

Confidence is deliberately kept separate from the score. A high score with low confidence means we do not yet have enough evidence, not that the person is weak. Where confidence falls below threshold, the result goes to a human reviewer instead of being published as a precise number.

Role weightings in use

Different roles weight the dimensions differently. These are starting configurations and are being validated against real job outcomes; they will change by version as we learn which signals actually predict performance.

brand designer

brand-designer-v1.0

  • Practical35%
  • Role competency20%
  • Portfolio evidence15%
  • Communication10%
  • AI fluency10%
  • Reputation/reliability10%

customer success

customer-success-v1.0

  • Communication25%
  • Role competency20%
  • Practical scenario20%
  • English15%
  • Reputation/reliability10%
  • Tools/AI fluency10%

full stack developer

full-stack-developer-v1.0

  • Practical code35%
  • Role competency25%
  • Debugging/security15%
  • Communication10%
  • AI coding fluency10%
  • Reputation5%

Where AI is used, and where it is not

AI assists with

Structuring an imported CV into a draft profile, which the talent must confirm.

Generating question variants from an approved competency blueprint.

Applying a written rubric to a text, voice or artifact response.

Transcribing audio and generating captions.

Turning an employer description into a structured role brief.

Phrasing a match explanation from features that already exist in the ranking data.

AI never

Decides the competency weights or the pass threshold at run time.

Publishes a score without schema validation and version metadata.

Issues a verification badge or removes access to opportunities on its own.

Adds a reason to a match explanation that is not present in the feature data.

Infers personality, honesty or emotion from a face or a voice.

Claims to determine whether a design or document was AI-generated.

Every evaluator run records the provider, model, prompt template version, rubric version and a confidence value, so any published score can be reconstructed and audited later.

Verification states

StateMeaningWhat employers see
Profile createdSelf-declared information, not yet checked.No badge.
Identity verifiedIdentity check passed.Identity badge only.
Assessment completedRequired assessments done, quality review may still be pending.Scores provisional and visible to the talent only.
Taskopedia VerifiedRole thresholds met, required checks passed, human QA complete.Verified badge with date and assessment version.
Verified, active evidenceRecent assessment or work evidence on file.Full badge.
Verification staleEvidence has aged beyond the freshness window.Badge annotated as reassessment due, or downgraded by policy.
Under reviewA material flag or appeal is open.Affected scores limited until resolved. This does not imply wrongdoing.

Common questions

Does AI decide whether I pass an assessment?
No. AI applies a written rubric to produce a per-item score, but the weights, the pass thresholds and the aggregation are held in versioned configuration and computed in code. Low-confidence and borderline results are routed to a human reviewer before anything is published, and no badge is issued without human quality review.
What does the confidence value next to a score mean?
Confidence is calculated separately from the score, as the product of how much of the assessment blueprint the evidence actually covers, how reliable the instrument is, and whether a human reviewed the result. A high score with low confidence means we do not yet have enough evidence, not that the person performed poorly.
Can I challenge a result I think is wrong?
Yes. Any score that materially affects your verification status or your access to opportunities can be appealed. An appeal creates a tracked case with a named reviewer who was not involved in the original decision. If a score is overridden, the original result is preserved alongside the new one with the rationale and policy basis recorded.
Do you assess personality, confidence or honesty from my video?
No. We do not use facial expression, emotion inference, honesty detection, attractiveness scoring or accent ranking. Communication is scored on observable criteria such as clarity, structure, professionalism, completeness and expectation setting, using the transcript and a written rubric.
Can I use AI in a practical assessment task?
Yes, and we assess how well you do it. AI fluency is a scored dimension covering whether you use AI effectively, verify its output and understand its limitations. What matters is that you disclose how AI was involved and that the resulting work is good. We do not claim to detect AI-generated work from the artifact alone.
Will my scores change if you update a rubric?
No. Every result is pinned to the assessment version it was taken under, and published assessment versions are immutable. Changing a rubric affects future sessions only. If you want a result under a newer version, you take a reassessment.
How long does verification stay valid?
Each role track has a freshness window. Once your most recent assessment or completed work falls outside it, your profile is marked as due for reassessment rather than presented to employers as current evidence.

Integrity and appeals

Assessments use randomised item pools, server-side timing and per-section limits. Where integrity signals are recorded, they are disclosed in advance, and they never automatically fail anyone. A flag routes a session to a human reviewer who looks at the actual evidence.

We do not use invasive webcam proctoring by default. Any future use would require a clear justification, explicit consent and legal review first.

Any score that affects verification or access to opportunities can be appealed. An appeal creates a tracked case with a named reviewer. If a score is overridden, the original result is preserved alongside the new one with the reviewer, rationale and policy basis recorded.