Taxonomy mismatch
A Skills Taxonomy Is a Translation Layer, Not a Universal Truth
O*NET, ESCO, Lightcast and SFIA describe work for different purposes. Treating any one of them as the workforce itself turns a mapping problem into an HR fact.
What the wrapper says
Skill taxonomies turn the ability described by skills and competencies into interoperable labels, without making the labels the ability.
The all-clear
O*NET, ESCO, Lightcast Open Skills, and SFIA are useful classification systems because each makes a different workforce decision more legible.
What is durable
The ability to perform a defined task or exercise sound judgment under stated conditions, independent of the vocabulary used to describe it.
The spreadsheet is called global_skills_normalised_FINAL_v7.xlsx.
Four tabs.
A locked column called standardized_skill.
And a note from somebody in procurement saying all duplicates have been removed.
They have not.
They have been renamed.
One record says network security.
Another says cybersecurity.
A third says information assurance.
A fourth is a Lightcast Open Skills identifier.
Somebody has helpfully mapped all four to an occupation, because a field named occupation_code was available and systems dislike empty cells.
Then the talent marketplace asks which people can do the work.
This is where a small clerical decision becomes a workforce fact.
A capability is not a taxonomy entry.
A capability is the ability to perform a defined task, or exercise sound judgment, under stated conditions.
An occupation is a grouping of work.
A taxonomy identifier is a label issued by an institution, for a purpose.
Related objects.
Not interchangeable ones, however urgently the dashboard requires a single filter.
The distinction is not decorative.
Lose it and you start treating a mapping decision as a discovery about your own people.
Every identifier arrives with an institutional purpose
O*NET-SOC is a United States occupational information architecture, aligned to the national Standard Occupational Classification.
That is not an attempt to describe the one true shape of human work.
It is an extremely useful way to organise occupational information for one country.
Its content model separates worker characteristics, worker requirements and occupational requirements, so an employer can reason from occupations to a consistent set of descriptors.
It does not make an O*NET occupation identical to a skill, a local job, or a person. Which matters rather a lot to the person whose title was typed into Workday during a reorganisation.
ESCO has a different job.
The European Commission’s multilingual linked-open-data classification carries 13,939 skill and knowledge concepts and 3,039 occupations, across 27 official European languages plus Ukrainian and Arabic.
Its skills are organised as knowledge, skills, language skills and transversal skills.
It uses SKOS and persistent, dereferenceable URIs.
And it explicitly relates skills to occupations as essential or optional.
That is a strong design for multilingual matching.
It is also a useful warning to anyone who believes every link means the same thing.
An essential relationship is not an optional relationship wearing a more confident font.
Lightcast Open Skills has another mandate again.
It draws on machine-learning and natural-language pipelines that ingest job postings, résumés and worker profiles across international labour markets.
More than 34,000 distinct skills.
A three-tier hierarchy.
Thirty-one major categories.
A 14-day release cycle.
It separates specialized skills, common skills and certifications.
This is exactly the sort of system that notices the market has acquired a new software framework before the annual job-architecture workshop has located its flip chart.
That speed is not a flaw.
That is its assignment.
SFIA is useful precisely because it is narrower.
Used in more than 100 countries, it focuses on digital, computing, software engineering and technology-enabled business work.
It crosses more than 100 professional IT skills with seven levels of responsibility, from “Follow” to “Set Strategy, Inspire, Mobilise.”
Those levels are defined through autonomy, influence, complexity, business skills and knowledge.
So SFIA can express a distinction an occupation label regularly conceals. Two people may hold the same technical capability while operating with very different responsibility and organisational influence.
The four systems overlap because work overlaps.
They do not agree, because they were not commissioned to agree.
The version number is doing real work
Classifications revise themselves.
Codes change while titles stay the same.
Titles change while codes stay the same.
Occupations merge, split, and get folded into residual categories.
So a code is not a timeless identity.
A title is not a timeless identity.
That may feel discourteous to your database schema.
It remains true.
Crosswalks between editions are also lossy in a specific direction.
A many-to-one mapping cannot be read backwards to recover which original record existed.
Aggregate at the wrong level and a real distinction disappears, in perfectly valid arithmetic.
None of that is bad taxonomy work.
That is what crosswalking means.
Which is why the useful rule is painfully unglamorous.
Keep the capability statement next to the label.
Record the source taxonomy, the version, the identifier and the local label. Then record the relationship to a canonical capability: exact, broader, narrower, related, partial or unmapped.
Preserve the mapping direction and the confidence.
Nobody gets a tote bag for this.
It prevents the report from being fiction.
A classification can be accurate and still be incomplete
Every one of these systems contains entries that announce their own limits.
A residual category exists so that records have somewhere to go.
It is a filing outcome.
It is not a secret capability profile shared by everyone inside it.
A new and emerging occupation may carry an excellent description of work that almost nobody in your organisation has actually done.
The description is helpful.
It still does not prove that an employee with a superficially similar local title performs that work, has that evidence, or operates at the same responsibility level.
Read either of those as a capability claim and the taxonomy has not misled you.
Your implementation has.
The taxonomies are genuinely good, which is the difficulty
None of this argues for abolishing O*NET, ESCO, Lightcast or SFIA and returning to a folder of résumés named resume_new_final_USETHIS.pdf.
O*NET provides an occupational information architecture with empirical descriptors.
ESCO’s multilingual URIs and essential-or-optional relationships make cross-language matching workable.
Lightcast observes labour-market language on a cadence that recognises the software market does not wait for a steering committee.
SFIA gives digital organisations a disciplined way to talk about responsibility as well as skill.
Each is useful classification infrastructure.
The failure begins when you take one system’s identifier and make it the canonical fact of your workforce.
A market signal becomes a performance claim.
A national occupation becomes a global role.
A responsibility level becomes a job title.
A residual bucket becomes a reason to deny somebody a project.
The taxonomy did not do that.
The implementation did.
AI makes the shortcut look intelligent
For years, inconsistent labels were a nuisance handled by analysts with context, patience and an alarming number of conditional formatting rules.
Then talent platforms began ingesting profiles, learning histories, project requirements and job descriptions at scale.
Take an enterprise skills marketplace that indexes verified competencies, project histories and learning aspirations against a central taxonomy of more than 15,000 organisational skill nodes.
In its first three years it unlocks more than 500,000 project hours across 70 countries.
A system operating at that scale needs structure.
It cannot convene a meeting every time data stewardship arrives as data governance.
This is the hinge.
AI can match strings at scale, and strings become persuasive very quickly once they appear in a recommendation.
But a machine that maps labels without their scope and provenance does not find equivalence.
It produces a very fast assertion of equivalence.
The result may exclude a worker from a project.
Or recommend unnecessary training.
Or report a skill gap created entirely by incompatible identifiers.
The model should be allowed to propose a relationship.
It should not be allowed to erase the fact that the relationship is partial.
Build the translation layer, then leave room for judgment
A workable workforce model keeps three records distinct.
The capability record describes what a person can do, and under which conditions.
The role record describes your work grouping, its scope, its expected outcomes.
The taxonomy record stores the issuer, version, identifier, label and source context.
Mappings sit between those records.
Not inside the person.
Each mapping needs a relationship type, a direction, a scope, a confidence and a review status.
A one-to-many mapping stays one-to-many.
A many-to-one mapping says what it loses.
An unmapped record stays unmapped, until there is evidence to resolve it.
And a title-only or residual record queues for review, rather than being promoted to an authoritative capability claim because it happened to be machine-readable.
There is a further limit worth naming.
Competency systems can become so eager to be precise that they atomise work into nonsense. The United Kingdom’s National Vocational Qualification system in the 1990s grew to more than 800 micro-competency elements and performance criteria per qualification.
The signatures multiplied.
Professional synthesis, ethical discernment and adaptive judgment did not become easier to see.
So the translation layer does not promise to measure every ounce of competence.
It makes clear what a label is evidence of, what it is not evidence of, and where a person has to decide.
A modest ambition.
Also considerably more useful than asking whether Cybersecurity Specialist and Cyber Security Ninja are the same thing, then writing “yes” into a production table because the demo is on Thursday.
What the record establishes
- Separate a capability statement from an occupation and from a taxonomy identifier.
- Store the source taxonomy and version with every imported label.
- Model one-to-many, many-to-one, partial, and unmapped mappings explicitly.
- Keep a canonical capability statement beside every local taxonomy label.
- Route residual, title-only, and unmapped categories to human review.
- Preserve mapping scope and confidence when AI uses taxonomy data for matching.
Asked in the review
- What is the difference between a skill, an occupation, and a taxonomy code?
- A skill is an ability to perform work or exercise judgment under stated conditions. An occupation is a role grouping that can require many skills. A taxonomy code is an identifier assigned by a particular classification system. O*NET-SOC 2019, for example, assigns occupational identifiers within a structure aligned to the 2018 Standard Occupational Classification system. The code describes a record in that system; it does not become the worker's capability.
- Why does every imported skills label need a source taxonomy and version?
- A label without provenance cannot be interpreted or maintained safely. O*NET-SOC 2019 records that 107 occupations change code from O*NET-SOC 2010, while 75 change title. A source taxonomy and version preserve the meaning that applied when the record arrived. They also make a later migration auditable instead of turning a changed code or revised label into an unexplained change in workforce capability.
- Can one taxonomy code map to more than one code in another system?
- Yes. Taxonomy mappings can be one-to-many, many-to-one, partial, or unavailable. In the O*NET-SOC 2019 crosswalk to the 2018 SOC, both Chief Executives and Chief Sustainability Officers map to the single 2018 SOC code 11-1011. Aggregating at the SOC level loses the distinction between those O*NET-SOC occupations. A mapping table therefore records direction, scope, and the fact that information is lost.
- What should sit beside a local taxonomy label in a skills system?
- A local label sits beside a canonical capability statement, not in place of one. The canonical statement describes the task, knowledge, judgment, conditions, or level of responsibility that matter to the organisation. The local label remains useful provenance: it can support reporting, labour-market analysis, or role design. Keeping both fields allows local classifications to change without silently redefining the underlying capability.
- What happens when a record lands in a residual or title-only category?
- A residual or title-only category requires review rather than automatic skill inference. O*NET lists 93 title-only occupations with no descriptor data, and describes 11-9039.00 Education Administrators, All Other as education administrators not listed separately. A residual category is a filing outcome, not evidence that its members share a capability profile. The record can be retained, but its uncertainty remains visible.
- Are O*NET, ESCO, Lightcast Open Skills, and SFIA still useful?
- Yes. Each system serves a useful but different purpose. O*NET-SOC 2019 organizes U.S. occupational information; ESCO 1.2 supports multilingual European skills and occupation data; Lightcast Open Skills captures labour-market language on a recurring release cycle; and SFIA profiles digital capability through skills and responsibility levels. Their usefulness is the reason an organisation should preserve their provenance rather than pretend they make the same claim.
- Why is ESCO not just a European version of O*NET?
- ESCO 1.2 is a multilingual linked-data classification with 13,939 skill and knowledge concepts and 3,039 occupations across 27 official European languages plus Ukrainian and Arabic. It uses SKOS and persistent URIs, and marks skill-to-occupation relationships as essential or optional. O*NET-SOC 2019 is a U.S. occupational information architecture aligned to the 2018 SOC. Similar subject matter does not make their identifiers interchangeable.
- Why does a market-facing taxonomy need different treatment from a public classification?
- Lightcast Open Skills is built from machine-learning and natural-language-processing pipelines that ingest job postings, résumés, and worker profiles across international labour markets. Its library contains more than 34,000 skills, organized under 31 major categories, and releases updates on a 14-day cycle. That makes it useful for detecting demand language, but its release cadence and source material are different from a public occupational classification's mandate.
- How should a system represent seniority in a technical capability model?
- A system represents responsibility separately from the skill name. SFIA uses a two-dimensional model of more than 100 professional IT skills and seven levels of responsibility. Its levels are defined through Autonomy, Influence, Complexity, Business Skills, and Knowledge. This structure shows why a label such as a technical skill is incomplete on its own: the same skill can be exercised with materially different scope and accountability.
- What should an AI matching system keep when it maps skills labels?
- An AI matching system keeps the source taxonomy, version, target taxonomy, mapping direction, relationship type, scope, and confidence with each match. A string match can be useful evidence, but it does not prove semantic equivalence. ESCO's essential and optional relationships, O*NET-SOC's many-to-one crosswalks, and SFIA's responsibility levels all show that the surrounding relationship changes what a label means in a decision.
- Why should unmatched categories go to a person for review?
- Unmatched categories reveal a boundary in the taxonomy, the mapping, or the incoming evidence. Automatic substitution can invent a capability claim that the source record does not support. This matters particularly when a framework reduces complex practice into small statements: the historical United Kingdom NVQ system grew beyond 800 micro-competency elements per qualification, illustrating how detailed labels can still fail to capture professional synthesis and judgment.