What we measure
We publish a multidimensional scorecard rather than a single mysterious number. Each dimension is scored separately, carries its own confidence value, and links to the evidence behind it.
| Dimension | What it means | Evidence |
|---|---|---|
| Role Competency | Knowledge and applied understanding of the target role. | Knowledge items and scenario tasks. |
| Practical | Quality of a real work simulation. | Work artifact scored against a rubric, plus reviewer check. |
| Communication | Written and voice client communication. | Situational responses and intro video transcript. |
| English | Functional workplace English. | Written, listening and speaking tasks. |
| AI Fluency | Effective, safe and verified AI-assisted workflow. | Scenario and practical exercise. |
| Portfolio Evidence | Strength and verifiability of prior work. | Source files, version history, client confirmation. |
| Reputation | Performance inside Taskopedia. | Reviews, completion, repeat work, disputes. |
| Reliability | Delivery behaviour, not personality inference. | On-time milestones, responsiveness, attendance where contracted. |
How a score is calculated
A dimension score is the weighted mean of the scored items within it. Confidence is calculated separately, as the product of how much of the blueprint the evidence actually covers, how reliable the assessment instrument is, and whether a human reviewed the result.
dimension_score = Σ(item_score × item_weight) / Σ(item_weight) confidence = evidence_coverage × assessment_reliability × review_factor published_score = round(dimension_score)
The weights are held in versioned configuration that is reviewed and published, never inside a prompt. A result is always pinned to the assessment version it was taken under, so changing a rubric never silently rewrites scores that already exist.
Confidence is deliberately kept separate from the score. A high score with low confidence means we do not yet have enough evidence, not that the person is weak. Where confidence falls below threshold, the result goes to a human reviewer instead of being published as a precise number.
Role weightings in use
Different roles weight the dimensions differently. These are starting configurations and are being validated against real job outcomes; they will change by version as we learn which signals actually predict performance.
brand designer
brand-designer-v1.0
- Practical35%
- Role competency20%
- Portfolio evidence15%
- Communication10%
- AI fluency10%
- Reputation/reliability10%
customer success
customer-success-v1.0
- Communication25%
- Role competency20%
- Practical scenario20%
- English15%
- Reputation/reliability10%
- Tools/AI fluency10%
full stack developer
full-stack-developer-v1.0
- Practical code35%
- Role competency25%
- Debugging/security15%
- Communication10%
- AI coding fluency10%
- Reputation5%
Where AI is used, and where it is not
AI assists with
Structuring an imported CV into a draft profile, which the talent must confirm.
Generating question variants from an approved competency blueprint.
Applying a written rubric to a text, voice or artifact response.
Transcribing audio and generating captions.
Turning an employer description into a structured role brief.
Phrasing a match explanation from features that already exist in the ranking data.
AI never
Decides the competency weights or the pass threshold at run time.
Publishes a score without schema validation and version metadata.
Issues a verification badge or removes access to opportunities on its own.
Adds a reason to a match explanation that is not present in the feature data.
Infers personality, honesty or emotion from a face or a voice.
Claims to determine whether a design or document was AI-generated.
Every evaluator run records the provider, model, prompt template version, rubric version and a confidence value, so any published score can be reconstructed and audited later.
Verification states
| State | Meaning | What employers see |
|---|---|---|
| Profile created | Self-declared information, not yet checked. | No badge. |
| Identity verified | Identity check passed. | Identity badge only. |
| Assessment completed | Required assessments done, quality review may still be pending. | Scores provisional and visible to the talent only. |
| Taskopedia Verified | Role thresholds met, required checks passed, human QA complete. | Verified badge with date and assessment version. |
| Verified, active evidence | Recent assessment or work evidence on file. | Full badge. |
| Verification stale | Evidence has aged beyond the freshness window. | Badge annotated as reassessment due, or downgraded by policy. |
| Under review | A material flag or appeal is open. | Affected scores limited until resolved. This does not imply wrongdoing. |
Common questions
- Does AI decide whether I pass an assessment?
- No. AI applies a written rubric to produce a per-item score, but the weights, the pass thresholds and the aggregation are held in versioned configuration and computed in code. Low-confidence and borderline results are routed to a human reviewer before anything is published, and no badge is issued without human quality review.
- What does the confidence value next to a score mean?
- Confidence is calculated separately from the score, as the product of how much of the assessment blueprint the evidence actually covers, how reliable the instrument is, and whether a human reviewed the result. A high score with low confidence means we do not yet have enough evidence, not that the person performed poorly.
- Can I challenge a result I think is wrong?
- Yes. Any score that materially affects your verification status or your access to opportunities can be appealed. An appeal creates a tracked case with a named reviewer who was not involved in the original decision. If a score is overridden, the original result is preserved alongside the new one with the rationale and policy basis recorded.
- Do you assess personality, confidence or honesty from my video?
- No. We do not use facial expression, emotion inference, honesty detection, attractiveness scoring or accent ranking. Communication is scored on observable criteria such as clarity, structure, professionalism, completeness and expectation setting, using the transcript and a written rubric.
- Can I use AI in a practical assessment task?
- Yes, and we assess how well you do it. AI fluency is a scored dimension covering whether you use AI effectively, verify its output and understand its limitations. What matters is that you disclose how AI was involved and that the resulting work is good. We do not claim to detect AI-generated work from the artifact alone.
- Will my scores change if you update a rubric?
- No. Every result is pinned to the assessment version it was taken under, and published assessment versions are immutable. Changing a rubric affects future sessions only. If you want a result under a newer version, you take a reassessment.
- How long does verification stay valid?
- Each role track has a freshness window. Once your most recent assessment or completed work falls outside it, your profile is marked as due for reassessment rather than presented to employers as current evidence.
Integrity and appeals
Assessments use randomised item pools, server-side timing and per-section limits. Where integrity signals are recorded, they are disclosed in advance, and they never automatically fail anyone. A flag routes a session to a human reviewer who looks at the actual evidence.
We do not use invasive webcam proctoring by default. Any future use would require a clear justification, explicit consent and legal review first.
Any score that affects verification or access to opportunities can be appealed. An appeal creates a tracked case with a named reviewer. If a score is overridden, the original result is preserved alongside the new one with the reviewer, rationale and policy basis recorded.