# Welcome

Welcome to Tabiya docs! This collection of documents provides an overview of our overall approach and our work.

<figure><picture><source srcset="/files/ZW6JwgYid9n34jNyHPI3" media="(prefers-color-scheme: dark)"><img src="/files/4nfimwmKDGJjDn5f0NpZ" alt="" width="375"></picture><figcaption></figcaption></figure>

Tabiya creates open-source software, models and standards to help tackle the global youth employment challenge. We foster research, coordination, and harmonization among partners that create learning and career pathways. Our vision is to unlock human capital and empower people in informal and formal labor markets.&#x20;

We are a non-profit organization that started at the [University of Oxford](https://www.oxfordmartin.ox.ac.uk/future-of-development/).

### Discover Tabiya's Work

<table data-view="cards"><thead><tr><th></th><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td><strong>Inclusive Livelihoods Taxonomy</strong></td><td>Make visible and usable the human capital of everyone in an economy</td><td></td><td><a href="/pages/QKz4o0shnBuLmjGYASlE">/pages/QKz4o0shnBuLmjGYASlE</a></td><td></td></tr><tr><td><strong>Compass by Tabiya</strong></td><td>Assist job-seekers in exploring and discovering their skills</td><td></td><td><a href="/pages/h41KegQPcMCNZoQrg4Cb">/pages/h41KegQPcMCNZoQrg4Cb</a></td><td></td></tr><tr><td><strong>Livelihoods Classifier</strong></td><td>Use natural language processing to automatically process unstructured text from jobseeker profiles and vacancies </td><td></td><td><a href="/pages/nxBNE5G1uQlxQ04gxuDL">/pages/nxBNE5G1uQlxQ04gxuDL</a></td><td></td></tr></tbody></table>

If you have any questions or comments, please do not hesitate to reach out via [email](mailto:hi@tabiya.tech). You can also create an [issue](https://github.com/tabiya-tech/docs/issues) in the [Github repository](https://github.com/tabiya-tech/docs/) that hosts this documentation.


# About Tabiya

In the next decade, about 1 billion young people will enter the labor force. We want to empower them with better data.

<figure><img src="/files/EMoFqqYJDcsyVG8rhTIQ" alt="" width="375"><figcaption><p>We aim to develop open-source technology and practical solutions to connect jobseekers with opportunities. The foundation of our work is a set of inclusive livelihoods taxonomies upon which partners can build applications that empower jobseekers.</p></figcaption></figure>

## Our Vision

Tabiya is a research organization that develops open-source technology and practical solutions to connect jobseekers with opportunities. Our vision is one of inclusive labor markets that empower people by recognizing their formal and informal skills.&#x20;

> Tabiya, from Arabic طبيعة “essence”, is a chess opening position that serves as the starting point for many possible subsequent moves. It alludes to the starting point from which jobseekers may navigate the many different paths through labor markets. It may also be translated as “talent”, alluding to the human capital that young people are looking to build, protect, and utilize.

Our objective is to create software, models, and standards that help governments, nonprofits, grassroots organizations, education and training providers, and employers in low- and middle-income countries (LMICs) harness the promise of data and AI to create more equitable and more efficient labor markets while navigating the global challenges of digitalization and decarbonization.&#x20;

### Our Guiding Principles

Our work is based on the following four guiding principles:

1. **Public goods instead of walled gardens:** Rich private sector ecosystems for labor market intermediation are emerging in low-and middle-income countries, but with little to no integration, coordination, and harmonization of data. LMIC governments and non-profits find it difficult to tap into private sector ecosystems for analytics or build public good solutions for their labor markets. Off-the-shelf labor market information systems and AI-based analytics are exclusively geared at high-income country labor markets and are prohibitively expensive for LMIC applications.
2. **Standardization to foster integration, coordination, and harmonization:** Different labor market actors need to speak the same “language” – they need to use a common reference terminology for the labor market. This reference terminology needs to describe the universe of jobs in an economy and the anatomy of each job – the mix of skills, competences, qualifications required for the job. Such frameworks either do not exist for LMIC labor markets or are frequently outdated. A standardized framework is a key requirement to harness data across actors for large-scale data analysis and AI-based applications. At the same time, a taxonomy that does not reflect livelihoods in LMICs risks exacerbating existing inequities and raises concerns about algorithmic fairness.
3. **Rigorous evidence and research on equity, efficiency, or algorithmic fairness:** There is an emerging research agenda on the promise and perils of technology and AI for labor market intermediation in high-income countries, but challenges in LMICs are not well understood.&#x20;
4. **A neutral platform for open discussion, exchange, and learning:** Coordination and knowledge exchange on the promise and perils of technology for labor markets among public and private employment service experts and policymakers from LMICs is lacking. With our roots at the University of Oxford, we provide a neutral platform to make recent advances in data and AI-based technologies more widely available while rigorously scrutinizing algorithmic biases.

## The Global Youth Employment Challenge

Youth employment is already a pressing global challenge, with about 75 million young people unemployed and 280 million people not in employment, education, or training. In the next 10 years, more than 1 billion young people will enter the labor force, 700 million in Africa alone.&#x20;

Over the next decades, digitalization and decarbonization will further transform livelihoods in low- and high-income countries alike. Jobs in various sectors of the economy might disappear or change dramatically while new economic opportunities will be created elsewhere. Investing into human capital and formulating policy responses to digitalization and decarbonization will require evidence, and targeted advice and recommendations to empower jobseekers in their journeys through these changing labor markets.&#x20;

Academic research has additionally identified a range of labor market frictions that undermine the efficient allocation of workers to jobs in LMICs. Millions of youths in low-income countries face a “slippery job ladder”: As they try to climb up to better-paying jobs, they frequently fall back into low-wage work, informal self-employment, or unemployment. This unproductive churn traps them in poverty and exacerbates inequalities. Lack of information is one explanation for why the ladder is so slippery – jobseekers often don’t know how to best use their skills, and firms don’t know where to find the best workers – and that improving information can improve jobseeker outcomes.&#x20;

Improved data and AI systems promise to tackle these challenges but remain inaccessible and unaffordable to stakeholders in LMICs. Proprietary commercial solutions dominate the market but raise concerns about limited interoperability and contractual lock-ins (“walled gardens”). Where these tools are used, they often exclude large parts of the population in the thriving and diverse informal sector. This exacerbates existing, often strongly gendered, inequities.

## Context: Improving labor markets with technology can only be a small puzzle piece of the global youth employment challenge&#x20;

Online job matching platforms have emerged as popular tool for connecting employers and job seekers in an increasingly digital world. And indeed, there is now robust evidence that alleviating search and matching frictions by improving the flow of information between employers and jobseekers can improve labour market outcomes.&#x20;

<details>

<summary>Why intervene in the labor market when jobseekers typically have strong incentives to find jobs, and firms have strong incentives to find the right employees?</summary>

Jobseekers typically have strong incentives to find jobs, and firms have strong incentives to find the right employees. As [Carranza and McKenzie *JEP* 2024](https://www.aeaweb.org/articles?id=10.1257/jep.38.1.221) show using Labor Force Survey data from a number of (largely middle-income) countries, public or private labor market intermediation services only help a tiny fraction of jobseekers find work.&#x20;

So what rationale is there to intervene in the free labor market and support job matching platforms and other innovative tech solutions that aim to reduce search and matching frictions?&#x20;

One important argument are equity concerns: We know that in many contexts, existing networks play a large role in who gets the few jobs that might be available. Underserved communities without these networks might fall short. Innovative, tech enabled solutions could, in theory, improve inclusion of these underserved groups if they provide information in a more equitable way. It remains, however, a question to what extend such platforms can truly improve economic inclusion for underserved groups. Not everyone has access to smartphones, the internet, or the necessary digital literacy to utilize these platforms effectively.&#x20;

A second argument might be that reducing information frictions can improve the reallocation of workers across sectors and space -- for example in contexts with rapidly changing demand in specific sectors. For example, the Ethiopian Investment Commission (EIC) – the Ethiopian government's investment promotion agency – worked with partners to set up a centralized sector-specific labor market information system to match a large supply of workers with specific jobs in the country's industrial parks.&#x20;

</details>

At the same time, it is important to note that information frictions only represent a very small part of the global youth employment challenge. While we believe that technology – if used well – can help overcome these frictions, they address merely the surface of deeper, structural issues – most importantly the lack of economic growth, insufficient job creation, and deeply rooted social and economic exclusion of specific groups. This broader economic context can render even the most efficient job matching platforms ineffective.&#x20;

Technology can thus only be part of a broader approach to solving global youth employment and promoting economic inclusion. It requires broad ecosystems of private sector, government, and civil society working together to tackle structural barriers on both sides of the market. Tabiya aims to contribute to these ecosystems with digital public goods that others can build on.


# The Global Youth Employment Challenge

Youth employment in low- and middle-income countries is in crisis, with rapid population growth outpacing job creation, leaving millions unemployed or trapped in informal, low-quality work.

Around the world, youth face disproportionately high unemployment and underemployment. The [global youth unemployment rate hovers around 13%](#user-content-fn-1)[^1] – roughly three times the adult rate (5.1%). In 2023, youth joblessness reached 20–25% in regions like the Middle East and North Africa, compared to about 9–10% in Sub-Saharan Africa and South Asia. But unemployment as an indicator understates the problem in poorer countries, where few can afford to remain jobless.&#x20;

An [estimated 500 million young people worldwide](#user-content-fn-2)[^2] were unemployed, underemployed, **or** working insecure jobs even before the COVID-19 shock. Currently, one of out five youth (\~260 million) are not in employment, education or training (NEET), [a number that had been steadily rising pre-COVID](#user-content-fn-3)[^3] and is [projected to continue rising](#user-content-fn-1)[^1]. The “youth employment crisis” cannot be explained through just a lack of jobs, but also a lack of **decent work**, **skills mismatches**, **youth expectation and aspiration gaps** and **structural barriers** that prevent young people in LMICs from securing decent livelihoods.

{% embed url="<https://ourworldindata.org/grapher/youth-not-in-education-employment-training?country=OWID_WRL~Sub-Saharan+Africa+(UN)~Central+and+Southern+Asia+(UN)~Europe+and+Northern+America+(UN)~Latin+America+and+the+Caribbean+(UN)~Least+Developed+Countries+(LDCs)~Northern+Africa+and+Western+Asia+(UN)&tab=chart&time=2010..latest>" fullWidth="false" %}

## Unpacking the crisis in low- and middle-income countries

LMICs face a perfect storm of challenges fueling youth unemployment – a booming working-age youth population outpacing formal job growth, education and training that often fail to match labor market needs, the dominance of informal employment with poor job quality, and structural barriers like restrictive regulations and urban-rural divides. These factors combine to make the school-to-work transition difficult for millions of young people. Below we break down these key challenges:

{% stepper %}
{% step %}

### Demographic Pressure

The “youth bulge” in developing countries means millions of new jobseekers each year, but economies are not creating enough formal jobs to absorb them. Over the next decade, about [1.1 billion young people](#user-content-fn-4)[^4] will enter the global labor force - a demographic wave concentrated in Asia and Afric&#x61;**.** In Sub-Saharan Africa alon&#x65;**,** [10-12 million youth will enter the labor market](#user-content-fn-5)[^5] across the region every year in the coming decade, yet only \~3 million new formal wage jobs are currently created each year. Such explosive growth in youth labor supply, without commensurate growth in stable jobs, creates intense competition for the limited opportunities available. Realizing the potential *demographic dividend* of these young populations will require economies [to generate **many more decent jobs**, and to do so quickly](#user-content-fn-6)[^6].&#x20;
{% endstep %}

{% step %}

### Skills Mismatch: the educated but unemployed

More than [57 out of 108 countries have a skills mismatch rate of over 50%](#user-content-fn-7)[^7] in their workforce. While this is primarily driven by young adult workers who are undereducated rather than overeducated, the latter has been steadily rising.&#x20;

In many countries, this [**skills mismatch** has been evident for years](#user-content-fn-8)[^8] – for example, large numbers of university graduates remain unemployed even as industries report unfilled job vacancies in technical and skilled trades. [In the early 2010s in Georgia](#user-content-fn-9)[^9], more than half of unemployed youth had a secondary school diploma and as many as 40% held a college degree. [A 2016 study ](#user-content-fn-10)[^10]covering comparing skills mismatches across [12 LMIC from 4 regions](#user-content-fn-11)[^11] found that on average only 52% workers were in jobs well-matched to their education. Of those mismatched, 36% were over-educated while about 12% were under-educated. Another [study using school-to-work transition surveys between 2012-13 from in 27 LMICs ](#user-content-fn-12)[^12]worldwide found that less than half of employees were considered well-matched. A [MENA study ](#user-content-fn-13)[^13]notes that, *college-educated youth unemployment coexists with shortages of IT, engineering, and other technical workers*, because many students graduate in fields with limited job prospects. This mismatch stems in part from curricula and training programs that are out of touch with labor market needs, as well as limited career guidance.&#x20;
{% endstep %}

{% step %}

### Informality and Job Quality

Because formal jobs are scarce, the vast majority of employed youth in LMICs end up in **informal employment**. This often means working on family farms, in petty trade, or self-employment in micro-businesses – typically without contracts, protections, or steady salaries. [In 2017, 77% of working youth were in informal jobs, compared to 58% of working adults.  ](#user-content-fn-14)[^14]These rates of informal employment among youth were much higher in developing and emerging economies.  [Young women face an even heavier burden](#user-content-fn-8)[^8]: in Sub-Saharan Africa and South Asia, 86–88% of young female workers are self-employed (mostly in informal work), far higher than for young men.&#x20;

<figure><img src="/files/GMouVrMij0A5kHa5PArq" alt=""><figcaption><p><strong>Youth informal employment rate (%), ages 15-24;</strong> Data Source: <a href="https://humancapital.worldbank.org/en/indicator/WB_HCP_EMP_NIFL_Y?geos=WLD">World Bank Human Capital Data Portal </a></p></figcaption></figure>
{% endstep %}

{% step %}

### Labor Mobility (or the lack thereof) and Structural Barriers

Labor mobility—both domestic and international—shapes youth unemployment in LMICs. Limited rural jobs push many young people to migrate to cities. While such rural-urban migration once improved individual prospects, it has been found to exacerbate urban joblessness – for example, [in 2018 youth unemployment rates were higher in cities](#user-content-fn-15)[^15] than in rural areas across Africa and the Middle East. Furthermore, **structural mobility restrictions** can lock youth out of opportunities. This includes poor transportation infrastructure, housing costs, or even formal restrictions on internal migration in some countries, which make it difficult for youth to relocate for jobs. International mobility offers potential opportunities but is [constrained by strict visa regimes, high costs, and information asymmetries](#user-content-fn-16)[^16], and vulnerable to[ increasing geo-economic fragmentation](#user-content-fn-17)[^17] - locking out many young people from working abroad.

In many LMICs, regulatory constraints and weak business climates impede job creation. Cumbersome business registration, high taxes or compliance costs, and rigid labor regulations (like high minimum wages or strict firing rules) may discourage firms from hiring formal employees – hitting inexperienced youth the hardest. For instance, if it’s costly or risky to take on a new worker, employers are less likely to give a chance to a first-time young jobseeker. Additionally, labor market institutions (public employment services, job information systems, etc.) are often underdeveloped, which exacerbates information gaps between youth and employers [(more on this in the following section)](#the-role-of-labor-market-intermediation).
{% endstep %}
{% endstepper %}

**Rapid technological change and subsequent labor demand shifts.** The [Future of Jobs Report 2025](#user-content-fn-17)[^17] projects that, in the next 5 years, 170 million jobs will be created while 92 million jobs are displaced, constituting a structural labor market churn of 22% of the 1.2 billion formal jobs in the dataset being studied. This pace of technological change may worsen the skills mismatches and access to opportunities mentioned above as traditional education systems struggle to keep up with rapidly evolving labor demand, a much larger concern for resource- and budget-constrained LMIC institutions than for those more proximate to the global technological frontier driving these changes. <br>

[^1]: International Labour Organization. (2024). *Global Employment Trends for Youth 2024 | International Labour Organization*. <https://www.ilo.org/publications/major-publications/global-employment-trends-youth-2024>

[^2]: *Toward Solutions for Youth Employment: A 2015 Baseline Report | International Labour Organization*. (2015, October 12). <https://www.ilo.org/publications/toward-solutions-youth-employment-2015-baseline-report>

[^3]: World Bank Open Data. (n.d.). *World Bank Open Data* \[Dataset]. Retrieved March 4, 2025, from <https://data.worldbank.org>

[^4]: Calculations using population projections from the United Nations' 2024 Revision of World Population Prospects

[^5]: African Development Bank Group. (2019). *Jobs for Youth in Africa, 2016-2025* \[Text]. African Development Bank Group. <https://www.afdb.org/en/documents/document/bank-group-strategy-for-jobs-for-youth-in-africa-2016-2025-89238>\
    \
    Between 2010-2017, total job creation of 9 million jobs per year kept pace with the working population growth, but only 2.6 million jobs out of these were formal wage jobs. &#x20;

[^6]: Akileswaran, K., Mazumdar, J., & Albertos, A. P. (n.d.). Youth Unemployment in the Developing World Is a Jobs Problem. *Stanford Social Innovation Review*. Retrieved March 4, 2025, from <https://ssir.org/articles/entry/youth_unemployment_in_the_developing_world_is_a_jobs_problem>

[^7]: *Empowering the workforce of tomorrow | UNICEF*. (2021, July 13). <https://www.unicef.org/reports/empowering-workforce-tomorrow>

[^8]: International Labour Organization. (2015, October 12). *Toward Solutions for Youth Employment: A 2015 Baseline Report*. <https://www.ilo.org/publications/toward-solutions-youth-employment-2015-baseline-report>

[^9]: Bank, W. (2013). Georgia: Skills Mismatch and Unemployment, Labor Market Challenges. *World Bank Publications - Reports*, Article 15985. <https://ideas.repec.org//p/wbk/wboper/15985.html>

[^10]: Handel, M. J., Valerio, A., & Sánchez Puerta, M. L. (2016). *Accounting for Mismatch in Low- and Middle-Income Countries: Measurement, Magnitudes, and Explanations*. Washington, DC:  World Bank. <https://doi.org/10.1596/978-1-4648-0908-8>

[^11]: Ghana, Kenya, China (Yunnan only), Lao PDR, Vietnam, Sri Lanka, Armenia, Georgia, Macedonia, Ukraine, Bolivia, Colombia

[^12]: Sparreboom, T., & Staneva, A. (2014). *Is education the solution to decent work for youth in developing economies? Identifying qualification mismatch from 28 school-to-work transition surveys*. International Labour Organization. <https://www.researchgate.net/publication/297760658_Is_education_the_solution_to_decent_work_for_youth_in_developing_economies_Identifying_qualification_mismatch_from_28_school-to-work_transition_surveys>

[^13]: MENA: Middle East and North Africa \
    \
    International Labour Organization. (2015, October 12). *Toward Solutions for Youth Employment: A 2015 Baseline Report*. <https://www.ilo.org/publications/toward-solutions-youth-employment-2015-baseline-report>

[^14]: International Labour Organization. (2017). *Global Employment Trends for Youth 2017: Paths to a better working future*. <https://www.ilo.org/publications/major-publications/global-employment-trends-youth-2017-paths-better-working-future>

[^15]: Kabbani, N. (2019). *Youth employment in the Middle East and North Africa: Revisiting and reframing the challenge*. Brookings Institution. <https://www.brookings.edu/articles/youth-employment-in-the-middle-east-and-north-africa-revisiting-and-reframing-the-challenge/>\
    \
    Using ILOSTAT Database, accessed June 30, 2018.

[^16]: Ratha, D. K., De, S., Eung Ju, K., Plaza, S., Seshan, G. K., Shaw, W., & Yameogo, N. D. (2019). *Leveraging Economic Migration for Development: A Briefing for the World Bank Board* \[Text/HTML]. World Bank. <https://documents.worldbank.org/en/publication/documents-reports/documentdetail/en/461021574155945177>

[^17]: World Economic Forum. (2025). *The Future of Jobs Report 2025*. <https://www.weforum.org/publications/the-future-of-jobs-report-2025/digest/>


# The Role of Labor Market Intermediation

By connecting the dots between jobseekers and employers, good intermediation reduces information asymmetries, helps align skills supply with demand, and lowers the frictions that keep youth out of jobs. This includes better data systems, job-matching platforms, career guidance, training aligned to market needs (often with results-based incentives), targeted support for young jobseekers, and digital innovations – including open-source tools – that make job markets more transparent and efficient.&#x20;

Below we delve into the different ways in which labor market intermediation addresses youth employment, followed by the promise of digital platforms and open source initiatives in shaping inclusive and accessible intermediation for the underserved.&#x20;

{% stepper %}
{% step %}

### Reduces Search Costs&#x20;

Labor market intermediation plays a crucial role in **connecting jobseekers with opportunities**, especially for inexperienced young entrants who lack professional networks. Improved intermediation can close this gap by aggregating job listings, validating candidates’ skills, and streamlining matches. These efforts reduce time and friction (“search costs”) in the hiring process in mainly two ways:&#x20;

* **By improving market transparency:** In many LMICs, information asymmetry is a big problem – youth often do not know where jobs are and employers struggle to identify reliable, qualified candidates. Labor market intermediation helps improve circulation of information in the market through on one hand, aggregating and actively publicizing vacancies, and on the other hand, by guiding jobseekers in properly positioning themselves based on firm/industry expectations (career counseling, CV guidance, interview coaching) &#x20;
* **By making information flows more targeted and efficient:** Even with the rise of online job portals, the [World Bank ](#user-content-fn-1)[^1]notes that obtaining the right kind of information remains a challenge for both job-seekers and employers. So now, modern job-matching platforms go beyond simple job boards: they can use advanced technology to pair candidates with suitable jobs in a more targeted way. For example, digital platforms that collect profiles of jobseekers (skills, education, and work preferences) and vacancies from employers, can algorithmically suggest good matches, alert youth to openings, and help employers find talent they might otherwise overlook.&#x20;
  {% endstep %}

{% step %}

### Bridging the Skills Gap

To address the skills mismatch, many countries have been pursuing demand-driven training and education reforms as part of labor market intermediation. Demand-driven approaches start by asking: *What skills are actually in demand by employers?* and then working backward to shape training programs accordingly. This often involves close partnerships between training providers and industry. For example, some successful youth employment programs in Latin America have combined classroom vocational training with on-the-job internships in sectors that are hiring – essentially guaranteeing that what youth learn is immediately relevant to available jobs. Evaluations of such programs find positive effects on employment and earnings (e.g. [Dominican Republic](#user-content-fn-2)[^2], Peru[^3], Colombia[^4], Argentina[^5]). Key features include teaching technical skills that local employers are seeking, as well as “soft” skills and work readiness, and then placing youth into apprenticeships or short-term jobs to gain experience.&#x20;
{% endstep %}

{% step %}

### Reducing Labor/Hiring Costs

On the demand side, employer incentives can encourage firms to hire youth whom they might otherwise consider “high risk” due to lack of experience. Other demand-side measures include apprenticeship incentives (offsetting the cost for companies to train novices) and reducing regulatory burdens for hiring youth (for instance, providing a tax break or easing strict labor rules for youth apprenticeships).
{% endstep %}
{% endstepper %}

[^1]: Solutions for Youth Employment. (2023). *The Use of Advanced Technology in Job Matching Platforms: Recent Examples from Public Agencies | Solutions For Youth Employment*. World Bank. <https://www.s4ye.org/node/4078>

[^2]: Card, D., Ibarrarán, P., Regalia, F., Rosas-Shady, D., & Soares, Y. (2011). The Labor Market Impacts of Youth Training in the Dominican Republic. *Journal of Labor Economics*, *29*(2), 267–300. <https://doi.org/10.1086/658090>\
    \
    Ibarrarán, P., Kluve, J., Ripani, L., & Shady, D. R. (2019). Experimental Evidence on the Long-Term Effects of a Youth Training Program. *ILR Review*, *72*(1), 185–222.

[^3]: Díaz, J. J., & Rosas-Shady, D. (2016). Impact Evaluation of the Job Youth Training Program Projoven. *IDB Publications*. <https://doi.org/10.18235/0011732>

[^4]: Attanasio, O., Kugler, A., & Meghir, C. (2011). Subsidizing Vocational Training for Disadvantaged Youth in Colombia: Evidence from a Randomized Trial. *American Economic Journal: Applied Economics*, *3*(3), 188–220. <https://doi.org/10.1257/app.3.3.188>

[^5]: Alzúa, M. laura, Cruces, G., & Lopez, C. (2016). Long-Run Effects of Youth Training Programs: Experimental Evidence from Argentina. *Economic Inquiry*, *54*(4), 1839–1859. <https://doi.org/10.1111/ecin.12348>


# Digital Platforms and AI in LMIC Labor Market Intermediation

## Digital Job Platforms Expanding Employment Access in LMICs

Digital job-matching platforms have been transforming how workers in low- and middle-income countries (LMICs) connect with jobs across formal, informal, and high-skilled sectors. Traditionally, most workers in developing countries found jobs through personal networks or by directly approaching employers, with only a tiny fraction using employment agencies​. This reliance on informal networks [limited one's opportunities to immediate circle and neighborhood](#user-content-fn-1)[^1]​. Online platforms help overcome this information gap by aggregating job listings and making them easily accessible. They reduce search frictions by centralizing job information, letting workers find opportunities across locations and sectors, and even offering online tools for skills testing and verification. These efficiencies led by digital platforms have extended employment access to groups often underserved in traditional labor markets.&#x20;

* In the informal sector, **mobile and web-based job exchanges** now connect low-income and unorganized workers with work opportunities at scale. Such platforms leverage **mobile technology requiring none to low internet access**, using simple text-based interfaces to connect workers with employers, and dramatically increasing the visibility of informal workers to potential employers, helping even those without formal credentials find jobs.&#x20;
* **Gig economy apps** (from ride-hailing to online freelancing marketplaces) allow workers to earn income on a task or contract basis. [Web-based gig work alone now engages ](#user-content-fn-2)[^2]an estimated 4.4% to 12.5% of workers worldwide (full- or part-time) Including location-based apps (like rideshare and delivery), up to **12% of the global labor market** may already be gig workers. In developing countries, these platforms are opening **unique avenues for youth, women, and rural populations** who have been left out of traditional job markets​.&#x20;
  * [**Online gig work is seen to support economic inclusion** ](#user-content-fn-3)[^3]by providing opportunities for young people, women, low-skilled workers, and those in areas with few local jobs​. In fact, most online gig workers are youth under 30 seeking to earn or learn new skills, and women have been found to participate in the online gig economy at higher rates.[ Analysis on 17 countries found](#user-content-fn-3)[^3] that 42% of online gig workers were women, while women's participation in the general labor market in those countries was only 31.8%. In societies where cultural norms limit women's mobility or confine them to domestic responsibilities such as childcare, online gig work provides a practical solution, enabling them to earn an income while managing their household duties.
* Digital platforms have also **expanded pathways for high-skilled professionals** in LMICs. Through online freelancing websites and remote work portals, skilled workers in developing countries can now access clients and jobs globally. This effectively **“exports” skilled labor** services from LMICs and brings in earnings. [The growth has been striking:](#user-content-fn-3)[^3] in Sub-Saharan Africa, job postings on a major online work platform jumped **130% between 2016 and 2020**, far outpacing the 14% growth seen in North America. By reducing geographic barriers, these platforms integrate LMIC talent into international markets, creating opportunities for software developers, designers, writers, and other professionals to secure contracts that were once out of reach.

## AI-Powered Enhancements in Job Matching and Skills Alignment

Artificial intelligence (AI) and data analytics  have exponentially enhanced capabilities of digital platforms in job-matching, workforce analytics, and skills alignment. Modern online employment systems increasingly deploy machine learning algorithms to match job seekers with vacancies far more effectively than basic keyword searches. Unlike traditional job boards that rely on one-to-one keyword matching, AI-driven platforms can interpret the context and meaning of job requirements and candidate profiles. [The result is a **bi-directional best-fit matching** process, ](#user-content-fn-4)[^4]where the platform suggests the best candidates for an employer *and* the best jobs for a candidate, often with a “match score”.

* Beyond matching, AI enables **advanced jobseeker analytics and insights** that were previously unattainable. [Using big data techniques](#user-content-fn-5)[^5], platforms can analyze profiles and employment histories to identify trends and predict outcomes. For example, Belgium’s public employment service (VDAB) uses [AI-based statistical profiling that uses “click data” functionality ](#user-content-fn-6)[^6]to job-seekers’ click-activity and behavior to predict the time jobseekers are unemployed. The Austrian PES also developed a statistical model that estimates a job seeker’s probability of short-term and long-term unemployment, allowing counselors to target support to those at highest risk of staying jobless.&#x20;
* **Gamification of psychometric assessments** has also picked up speed in recent years, including among public employment services. In India, the National Skill Development Corporation (NSDC) partnered with KnackApp to develop a candidate profiling mechanism (skills, traits, and entrepreneurship potential) through cognitive games. This is used to guide students to career opportunities and jobs best suited to their interests.
* AI tools are also being used to **forecast labor demand** – [France’s “La Bonne Boîte”](https://labonneboite.francetravail.fr/questions-frequentes) uses a predictive algorithm to analyze recruitments from the past 12 months, to predict those for the next 3 months. This data enables jobseekers to identify a shortlist of companies 'with high hiring potential' to help target unsolicited applications. These kinds of analytics improve decision-making for both workers and policymakers: jobseekers get data-driven guidance (for example, which industries are growing or which skills are in demand), while governments obtain real-time labor market intelligence to design better training and employment programs.
* AI-driven career assistants are enhancing the **personalization of guidance and training** for workers on these platforms. Intelligent career assistants or “job coach” chatbots analyze user profiles, labor market data, and hiring trends to interactively provide tailored job recommendations, upskilling suggestions, and career coaching. However, the jury is still out on the effectiveness of these Ai-enabled career guidance tools. [An evaluation of ](#user-content-fn-7)[^7][Bob Emploi, an online job-search assistance website](#user-content-fn-8)[^8], found null effects across the board on key search and employment outcomes.

## The Shift Towards Skills and Competency-Based Profiling&#x20;

Crucially, AI is helping shift digital employment platforms toward **skills-based matching and alignment**. Traditional recruitment focuses heavily on formal qualifications and job titles, which can overlook candidates who have the right skills but non-linear backgrounds. AI allows platforms to parse rich data on hard and soft skills from resumes, online profiles, behavioral assessment and match those to job requirements​.&#x20;

**Advanced job platforms now often include a competency-based matching component:** rather than filtering candidates by degree or past job titles alone, the algorithm considers the full spectrum of technical skills, transferable skills, and even aptitudes. This holistic approach means a job seeker’s coding, language, or teamwork skills (even if self-taught or gained informally) can be recognized and matched to open positions, widening opportunities. It also helps employers discover talent that might be hidden in non-traditional resumes.&#x20;

**Many countries have have expanded traditional labor taxonomies and frameworks to include skills and competencies**. The European Commission's European Skills, Competences, Qualifications and Occupations (ESCO)[ ](#user-content-fn-9)[^9]now defines nearly 13,500 skills mapped to the ILO's pre-existing [ISCO ](#user-content-fn-10)[^10]occupational pillars. Frameworks like ESCO and O\*Net, its US-equivalent, provide a holistic mapping of occupations and skills, but have been challenging to localize and apply to more developing/emerging contexts - giving rise to [alternative methodological explorations](#user-content-fn-11)[^11] of leveraging big data from online job vacancies and applicants' profiles in combination with natural language processing (NLP) to extract information on skills and create local taxonomies from scratch.&#x20;

[^1]: McKenzie, D., & Carranza, E. (2023). *Government policy efforts on job search and intermediation: What works and what should be done better?* World Bank Blogs. <https://blogs.worldbank.org/en/impactevaluations/government-policy-efforts-job-search-and-intermediation-what-works-and-what>

[^2]: Alzate, D. (n.d.). *The Effects of Regulating Platform-based Work on Employment Outcomes: A Review of the Empirical Evidence* \[Text/HTML]. Retrieved March 7, 2025, from <https://documents.worldbank.org/en/publication/documents-reports/documentdetail/en/099100124224513583>

[^3]: Datta, N., Rong, C., Singh, S., Stinshoff, C., Iacob, N., Nigatu, N. S., Nxumalo, M., & Klimaviciute, L. (2023). *Working Without Borders: The Promise and Peril of Online Gig Work*. Washington, DC: World Bank. <https://doi.org/10.1596/40066>

[^4]: Solutions for Youth Employment. (2023). *The Use of Advanced Technology in Job Matching Platforms: Recent Examples from Public Agencies | Solutions For Youth Employment*. World Bank. <https://www.s4ye.org/node/4078>

[^5]: Desiere, S., Langenbucher, K., & Struyven, L. (2019). *Statistical profiling in public employment services: An international comparison* (OECD Social, Employment and Migration Working Papers 224; OECD Social, Employment and Migration Working Papers, Vol. 224). <https://doi.org/10.1787/b5e5f16e-en>

[^6]: Ernst, S., Mueller, A. I., & Spinnewijn, J. (2024). Risk Scores for Long-Term Unemployment and the Assignment to Job Search Counseling. *AEA Papers and Proceedings*, *114*, 572–576. <https://doi.org/10.1257/pandp.20241092>

[^7]: Ben Dhia, A., Crépon, B., Mbih, E., Paul-Delvaux, L., Picard, B., & Pons, V. (2022). *Can a Website Bring Unemployment Down? Experimental Evidence from France* (Working Paper 29914). National Bureau of Economic Research. <https://doi.org/10.3386/w29914>

[^8]: Bob Emploi endeavors to assist and motivate job seekers by providing personalized and data-informed advice on sectors and locations to target; o ering step-by-step planning assistance; sending regular reminders and encouragement messages; and providing general tips, such as how to behave during a job interview.

[^9]:

[^10]: International Standard Classification of Occupations

[^11]: Escudero, V., Liepmann, H., Boschetti Adamczyk, W., Boehmer, S., Delaporte, I., & International Labour Organization,. (2025). *Developing a new method to uncover skills trends in emerging economies using online data and NLP techniques*. ILO. <https://doi.org/10.54394/HQQX3200>


# Regional Spotlight: the Sub-Saharan African Labor Market

## Labour Market Frictions in Sub-Saharan Africa

One lever for addressing unemployment is helping firms to increase their hiring of workers from the large pool of available candidates. However, firms are often reluctant to expand hiring, and cite an inability to find suitable candidates. The World Bank reports that about 23% of firms cite workforce skills as a significant constraint to their operations. In some African and Latin American countries, this share rises to 40–60% (World Bank 2023). This is surprising when firms have large applicant pools available to them. Increasingly, academic research is showing that both firms and workers lack good information required to match worker skills to suitable jobs. This leads to workers searching for and applying for jobs inefficiently and firms being unable to find or observe the skills they need when hiring. These frictions reduce firm willingness to hire and reduce the wage offers firms are willing to make because of uncertainty of the quality of hired workers.&#x20;

Features common to many Sub-Saharan Africa labour markets are suggestive of large frictions that may reduce hiring demand and misdirect job search. Large shares of jobseekers have limited formal work experience and very similar education qualifications. Young jobseekers are particularly likely to have few experiences they can use to demonstrate skills. These constraints limit the ability for jobseekers to signal competency to firms and limit the ability for firms to compare applicants and judge who is best suited for the role. Moreover, job search and migration costs are high, especially relative to incomes. As summarised in Caria & Orkin (2024), studies from Ethiopia, Jordan, South Africa and Uganda find that job search expenses among active jobseekers are at least 16% of total jobseeker expenditure. This reduces the ability for job seekers to search widely, find the jobs that demand their particular abilities and experiences and gain information about the nature of jobs available in the market. Finally, a number of studies document a high prevalence of inaccurate beliefs among jobseekers (Hensel et al. 2024, Abebe et al., 2021). This suggests job seekers lack information about their skills, comparative advantage and prospective earnings.

Spatial frictions exacerbate difficulties for the matching of firms and workers. Studies that document high job search costs in Sub-Saharan Africa show that reducing travel costs by providing for example, transport subsidies or improving travel quality (Donald & Grosset, 2022) can increase job search intensity in the short term (Franklin et al. 2015, Abebe et al., 2021) but cannot entirely overcome frictions arising from weak signalling ability, misdirected search and biased beliefs. Intuitively, this is because searching for more jobs may not translate into matches if effort is focused on applying for jobs that applicants are not well suited for. Consequently, job search subsidies tend to fail to translate increased search into improved employment outcomes.<br>

Overcoming spatial Job fairs that bring together jobseekers with firms looking to recruit (Abebe et al., 2023) have been shown to provide information to both parties, leading to updated, more accurate, beliefs about the job market and pool of candidates. Online job search platforms can play a similar role in easing some costs of search and providing information about prospective employers and jobs for jobseekers (Wheeler et al., 2022., Field et al., 2023, Jones & Sen, 2022). However, a number of studies document no overall effect on employment from encouraging the use of online job platforms, and in some settings, a replication of off-platform negative bias towards socially marginalised groups (Chakravorty et al. 2023, Afridi et al 2023).

## Bibliography:&#x20;

&#x20;Abebe, G., Caria, A.S., Fafchamps, M., Falco, P., Franklin, S. and Quinn, S., 2021. Anonymity or distance? Job search and labour market exclusion in a growing African city. The Review of Economic Studies, 88(3), pp.1279-1310.

Abebe, G., Caria, S.A., Fafchamps, M., Falco, P., Franklin, S., Quinn, S. and Shilpi, F.J., 2023. Matching frictions and distorted beliefs: Evidence from a job fair experiment (No. 958). working paper.

Abel, M., Burger, R., Carranza, E. and Piraino, P., 2019. Bridging the intention-behavior gap? The effect of plan-making prompts on job search and employment. American Economic Journal: Applied Economics, 11(2), pp.284-301.

Abel, M., Burger, R. and Piraino, P., 2020. The value of reference letters: Experimental Evidence from South Africa. American Economic Journal: Applied Economics, 12(3), pp.40-71.

Afridi, F., Dhillon, A., Roy, S. and Sangwan, N., 2023. Social Networks, Gender Norms and Labor Supply: Experimental Evidence Using a Job Search Platform (No. 677). Competitive Advantage in the Global Economy (CAGE).

Banerjee, A. and Sequeira, S., 2023. Learning by searching: Spatial mismatches and imperfect information in Southern labor markets. Journal of Development Economics, 164, p.103111.

Bertrand, M. and Crépon, B., 2021. Teaching labor laws: Evidence from a randomized control trial in South Africa. American Economic Journal: Applied Economics, 13(4), pp.125-149.

Bhorat, H., Köhler, T. and de Villiers, D. (2023). Can Cash Transfers to the Unemployed Support Economic Activity? Evidence from South Africa. Development Policy Research Unit Working Paper 202301. DPRU, University of Cape Town.

Carranza, E., Garlick, R., Orkin, K. and Rankin, N., 2022. Job search and hiring with limited information about workseekers’ skills. American Economic Review, 112(11), pp.3547-3583.

Chakravorty, B., Bhatiya, A.Y., Imbert, C., Lohnert, M., Panda, P. and Rathelot, R., 2023. Impact of the COVID-19 crisis on India’s rural youth: Evidence from a panel survey and an experiment. World Development, 168, p.106242.

Fields, G.S., 2011. Labor market analysis for developing countries. Labour economics, 18, pp.S16-S22.

Franklin, S., 2015. Location, search costs and youth unemployment: A randomized trial of transport subsidies in Ethiopia.

Jones, S. and Sen, K., 2022. Labour market effects of digital matching platforms: Experimental evidence from sub-Saharan Africa.

Mudiriza, G., De Lannoy, A. (2023). Profile of young people not in employment, education or training (NEET) aged 15-24 years in South Africa: an annual update. Cape Town: Southern Africa Labour and Development Research Unit, University of Cape Town. (SALDRU Working Paper Number 298).

Wheeler, L., Garlick, R., Johnson, E., Shaw, P. and Gargano, M., 2022. LinkedIn (to) job opportunities: Experimental evidence from job readiness training. American Economic Journal: Applied Economics, 14(2), pp.101-125.\
Hensel, L., [Tekleselassie](https://cssh.northeastern.edu/faculty/tsegay-tekleselassie/), T., [Isphording](https://sites.google.com/view/ingoeisphording/about-me), I.,  [Radbruch](https://sites.google.com/site/jonasradbruch01/), J. & [Witte](http://www.marcwitte.com/home), M. 2024. Demand for Feedback and Job Search. Working Paper.


# Open-Source Tech for Labor Markets

## The Problem with Proprietary Systems and Vendor Lock-In

Across the world, important government information systems (for employment, education, etc.) have historically been implemented by a handful of proprietary software vendors, leading to high costs and vendor lock-in for public agencies​. This [proprietary control can stifle innovation and limit adaptability](#user-content-fn-1)[^1] of platforms to local needs – not to mention straining the budgets of resource-constrained LMIC governments.&#x20;

## The Benefits of Open-Source Solutions

**By contrast, adopting open-source software (OSS) solutions allows countries to avoid being locked into a single provider and instead foster local ownership and customization.** According to the World Bank, [implementations in several domains have been “dominated by a few IT vendors” ](#user-content-fn-1)[^1]with significant switching costs, but OSS offers a way to “facilitate efficiency, robustness, security, and interoperability” in these systems. In other words, when the code is open and reusable, each country or organization doesn’t need to reinvent the wheel or pay steep licensing fees for every new project. They can build on proven platforms, share improvements, and focus resources on reaching underserved populations rather than on proprietary licensing.

### Democratizing Access to Labor Market Tools

Open-source digital tools are also key to labor market inclusion - both in terms of drawing contributions from and enabling a wide range of actors and in reaching marginalized population segments.

* An open platform can be modified to accommodate multiple languages, literacy levels, and accessibility needs – critical for reaching rural and marginalized communities in LMICs. Open solutions can be designed to work in low-bandwidth environments or integrate with basic mobile phones, ensuring that even jobseekers without a smartphone or constant internet can access opportunities. This is particularly important with nearly 3 billion people in the world still offline as of 2022.
* Because the source code is open, solutions can be tailored to cope with irregular power supply, intermittent internet, or older devices commonly found in developing regions. For example, a lightweight version of a job portal could be developed to consume minimal data, or an offline mode could allow users to browse job info without continuous connectivity. By contrast, a proprietary system might not prioritize these LMIC-specific needs if they conflict with the vendor's global business model.

### Building Collaborative Ecosystems through Open-Source

Open-source digital tools  encourage broader partnership across the public and private sector: local tech communities or social enterprises can contribute modules (for example, an SMS-based interface for areas with low internet coverage) without needing permission from a vendor. By being free and openly available, such platforms lower barriers for small organizations or even community groups to deploy job-matching services in low-income areas.

Governments can mandate open data sharing from these systems (while protecting privacy), enabling researchers, start-ups, or nonprofits to build complementary services – from commute planning for jobseekers to analytics identifying skill gaps in the workforce. In Kenya and Nigeria, for instance, open data on job vacancies and skills demand has allowed innovators to develop third-party apps that target specific communities and industries.

### Ensuring Equity and Accountability

Another advantage of open-source labor platforms is transparency and trust. With algorithms open to scrutiny, there is greater accountability in how job matches are made or how data is used, which can protect against biases or unfair practices. It also means critical labor market data can remain in the public realm.

### Fostering Global Innovation and Knowledge Sharing

OSS as a public good drives opportunities for cross-country learning and collaborations. India’s Aadhaar identification system - fully built on open-source tech - has inspired development of initiatives like the [Modular Open Source Identity Platform](https://mosip.io/) (MOSIP) - which, [as of 2023](#user-content-fn-2)[^2], had scaled across 6 countries with ongoing pilots in 5, creating a network effect of better practices. &#x20;

### Open-Source: A Developmental Imperative?

Making digital labor tools open-source prevents proprietary gatekeeping of job information, empowers local innovation, and ensures that the benefits of digital intermediation extend to the hardest-to-reach populations. This ethos of openness means a nonprofit in Africa or a government in South Asia could take a similar tool, adapt it for their context (say, training the algorithm on local labor data), and deploy it for their citizens without starting from scratch. The result is a more level playing field where all countries, regardless of income level, can harness cutting-edge digital intermediation for their labor force.

[^1]: World Bank. (2019). *Open Source for Global Public Goods*. <https://openknowledge.worldbank.org/entities/publication/c9b2358e-16e0-5144-83fe-e4abe719f073>

[^2]: Hariharan, V. (2024, February 13). *Lessons Learned: Reflecting on MOSIP’s Journey to Scale*. Digital Impact Alliance. <https://dial.global/research/lessons-learned-reflecting-mosips-journey-scale/>


# Inclusive Livelihoods Taxonomy

We aim to make visible and usable the human capital of everyone in an economy. This page outlines our foundational approach to building a taxonomy that covers the full spectrum of economic activities.

## Understanding Human Capital

Human capital—the collective skills, knowledge, and experiences of individuals—is a key driver of economic productivity and income potential. The extent of an individual's human capital often determines their earning ability and is deeply tied to human agency, or the capacity to make independent decisions and shape one's life. Recognizing and valuing all forms of human capital is essential for creating inclusive economic opportunities.

The motivation for our work is described well by a quote from one our key partners, [Harambee Youth Employment Accelerator](https://www.harambee.co.za):

> *<mark style="color:$info;">Young people in South Africa often lack the resources, networks, education and work experiences needed to be considered for formal employment. But, in the past 12 years, our work at Harambee has taught us that young people have the potential to perform in these jobs if we give them a chance! What if a young person was better able to identify and articulate the skills they have gained outside of the formal economy? What if they could signal skills gained from unpaid work?</mark>*
>
> ***Harambee Youth Employment Accelerator, South Africa***

## Problem: Traditional Frameworks Overlook Parts of the Economy

Organizations supporting young people in accessing economic opportunities often use structured frameworks to define available opportunities and the skills, competencies, and qualifications required. However, most widely used labor market taxonomies fail to capture the diverse economic contributions of individuals. These frameworks, such as the **System of National Accounts (SNA)**—an international standard for measuring national economic activity—primarily account for formal, market-based transactions.

<details>

<summary><mark style="color:blue;">The System of National Accounts: What does it include, what does it exclude?</mark></summary>

The [System of National Accounts (SNA)](https://unstats.un.org/unsd/nationalaccount/sna.asp) serves as an international statistical standard for the measurement of economic activities. This methodological framework, employed by various countries around the globe, guides the production, interpretation, and use of internationally comparable economic statistics. The SNA's existence is predicated on the need to standardize and simplify the complex nature of economic transactions. It functions as an economic map, describing the interconnections between different economic actors (households, businesses, government), their activities (consumption, production, investment), and the overall performance of an economy.&#x20;

One key element in the SNA is the concept of the "production boundary." The production boundary delineates the transactions that are accounted for in the calculation of Gross Domestic Product (GDP) and other key economic indicators.&#x20;

</details>

However, many forms of human capital investment and productivity remain outside these traditional boundaries, including household labor, volunteer work, and informal employment. These activities play a crucial role in economic development and individual livelihoods yet remain undervalued because they do not generate direct monetary compensation. We refer to these overlooked activities as the **"unseen economy"**.

<figure><img src="/files/gu35IQIcbdaOg5jtadCr" alt=""><figcaption><p>Partners that help young people find jobs typically use a taxonomy to describe the universe of jobs in an economy and the anatomy of each job. Such a “map” of the labor market can help young people find escalators and elevators to better livelihoods. If the map is incomplete or inaccurate, it may limit young people’s ability to find the right jobs. <strong>We aim to expand the map to include the full range of diverse livelihoods that young people pursue, including those in the "unseen economy".</strong></p></figcaption></figure>

### **Recognizing the "Unseen Economy"**

The "unseen" parts of the economy must be acknowledged to ensure a comprehensive understanding of human capital. **Seen** economic activities typically involve paid labor, whether formal or informal. **Unseen** activities, by contrast, include unpaid but productive work—such as caring for family members or performing household tasks—that could, in theory, be outsourced for payment. Notably, this definition excludes leisure, as one cannot pay someone to engage in leisure on their behalf.

Recognizing and accounting for the human capital in the unseen parts of the economy is vital for accurately representing people's contributions to the economy, and providing a basis for policies that protect and support all forms of work, thereby enhancing individual agency.

<details>

<summary><mark style="color:blue;">Not making visible and usable human capital from the "unseen economy" can limit human agency in at least four ways</mark></summary>

1. **Undervalued Skills & Experience:** Many skills gained in unpaid or informal work—such as budgeting, logistics, and negotiation—are highly transferable. However, when these activities are not recognized as productive, individuals (especially women, who disproportionately handle unpaid labor) face challenges transitioning into paid employment or receiving fair compensation. Contreras et al. (2024) show that listing household activities in surveys increases reported labor force participation, suggesting that clearer recognition of skills helps job seekers articulate their experience.
2. **Limited Economic Opportunities:** The lack of recognition and compensation for unpaid work leaves many without financial resources to invest in education, entrepreneurship, or career growth. This perpetuates income inequality, poverty, and restricted social mobility, limiting individual agency.
3. **Policy Blind Spots:** When unpaid and informal work is excluded from economic metrics, it is often overlooked in policy decisions. As a result, support structures like social security, healthcare, and labor protections typically focus on paid workers, leaving many without essential safety nets.
4. **Reinforcement of Gender Inequalities:** Women globally bear a disproportionate share of unpaid work, such as caregiving and household labor. The failure to value this work limits their participation in paid employment, skill development, and broader economic opportunities, reinforcing systemic gender disparities.&#x20;

</details>

## Our approach: Expanding the Map

Our mission is **to make visible and usable the human capital of everyone in an economy**.  This involves two related efforts:

1. **Making visible** all economic activity and the skills, knowledge, and experience gained through these activities, particularly those in traditionally "unseen" sectors, such as informal and unpaid work.
2. **Making usable** the skills, knowledge, and experience gained through these activities by integrating them into conventional labor market frameworks.

The first has long been explored in social science and economic research, a base that we continue to utilize and build upon through our own research. The second objective of making these unseen skills usable is accomplished by expanding the existing map of the labor market: **creating an inclusive reference taxonomy that partners can adapt and build upon.**&#x20;

<details>

<summary><mark style="color:blue;">What is a reference taxonomy and why is it important?</mark> </summary>

A well-structured taxonomy provides a common language to categorize, interpret, and link various labor market data points. It serves as a **map of the labor market**, outlining the full spectrum of jobs in an economy and the competencies, skills, and qualifications required for each. These taxonomies serve as a foundation for governments, nonprofits, and other stakeholders to offer more effective career guidance and employment services. Below, we outline various use cases for this taxonomy.&#x20;

* **Matching:** A reference taxonomy bridges labor market supply and demand by standardizing and categorizing skills and qualifications. It helps match job seekers to roles by aligning employer requirements with candidates' capabilities.
* **Career Guidance & Skill Development:**  A reference taxonomy can be used for tailored guidance to jobseekers, highlighting required skills for specific roles and suggesting alternatives based on transferable competencies.
* **Data Analysis & Insights:** Standardized classification of skills, qualifications, and job titles enables meaningful labor market analysis, revealing trends, in-demand skills, and industry hiring patterns.
* **Policy & Research:** A unified classification system supports policymakers and researchers by enabling consistent labor market comparisons across industries, regions, and time, aiding workforce planning and education strategies.

</details>

With Tabiya's **Inclusive Livelihoods Taxonomy**, we aim provide a more inclusive map of the labor market – one that includes activities from the "unseen economy." A more inclusive map of the labor market will allow more inclusive matching, the identification of more diverse career and skill development pathways, and richer data analysis. [A more detailed description of our methodology can be found here](/our-tech-stack/inclusive-livelihoods-taxonomy/methodology) and our reference taxonomy can be accessed, adapted, and updated through our [Open Taxonomy Platform](/our-tech-stack/inclusive-livelihoods-taxonomy/open-taxonomy-platform).


# Methodology

Creating an inclusive livelihoods taxonomy involves assigning human capital and skills to activities that are usually unseen, which comes with challenges.

Tabiya's work on the inclusive livelihoods taxonomy can be divided in two streams:&#x20;

1. **Adaptation of the Seen Economy.** Making sure that the "seen" part of the pre-existing taxonomies such as ISCO and ESCO adequately fit local contexts. This work is highly specific to country contexts, and [our rigorously tested approach in South Africa is described in details here. ](/tabiya-south-africa/seen-economy)
2. **Making Visible the Unseen Economy.** Broadening existing taxonomies so that they include the unseen part of the economy, namely the activities that are typically not considered as productive and the skills that are associated with them.&#x20;

In this section we detail our approach to the second workstream - our work expanding the map of the labor market.&#x20;

## Motivation of the Methodology

<mark style="background-color:yellow;">**The challenge is the following: Tabiya aims at creating an inclusive taxonomy of livelihoods, while ensuring that this taxonomy is compatible with existing ones such as ISCO or ESCO.**</mark> This allows us to ensure that the developed taxonomy can be used by institutions that rely on globally standardized and widely-used taxonomies. We chose to base our work on the European Skills, Competencies, Qualifications and Occupations (ESCO) taxonomy ([reasons detailed here](/our-tech-stack/inclusive-livelihoods-taxonomy/why-esco)). &#x20;

The first step, naturally, was to evaluate the inclusiveness of the existing ESCO taxonomy. The intellectual underpinning of this work builds upon the "Counting Women's Work" literature, that aims at assigning a monetary value to the tasks done by women, especially in their households. In Tabiya's work, we aim to highlight the human capital gained from these tasks more than assigning monetary values. Additionally, our work covers all job-seekers, not just women, although depending on local contexts they may *de facto* represent the majority of unseen job-seekers.&#x20;

{% hint style="success" %}
**Counting Women's Work**: The necessity to better understand sex inequalities in the labor market led to the finding that while both men and women work, their work is valued differently, and this differential valuation yields a perceived differential “productive characteristic endowment” amongst sexes that, in turn, drives sex wage disparities for the African case. In particular, men generally perform paid labor market activities, while women perform both paid labor market and unpaid home production activities ([Dinkelman & Ngai, 2022](https://www.aeaweb.org/articles?id=10.1257/jep.36.1.57)). Hence, a literature that seeks to count women’s work has emerged. This literature relies on time-use data to attribute an equivalent labor market wage to home production activities done by women ([Samarasinghe, 1997](https://books.google.co.uk/books?hl=en\&lr=\&id=TGdd8mY3heQC\&oi=fnd\&pg=PA129\&dq=counting+women%27s+work\&ots=7mx1tLfjvK\&sig=dx9mmtD3ZJbOLGHTURnOBr8nlJI\&redir_esc=y#v=onepage\&q=counting%20women's%20work\&f=false); [Hoskyns and Rai, 2007](https://www.tandfonline.com/doi/full/10.1080/13563460701485268?casa_token=V92z2s5OCJ4AAAAA%3Ac8Ty5sBz0GYyaJ8RBVbAvmNa7FAycOeVS4sGuxvapqLar8hHoY15BjJO-mAIoBr3-7ut8VHJtZ0VkA); [Donehower, 2018](https://static1.squarespace.com/static/5994a30fe4fcb5d90b6fbeab/t/5bac023d4785d3a47239adb2/1537999437377/CWW+WP4.pdf) and [Abrigo & Francisco-Abrigo, 2019](https://www.econstor.eu/handle/10419/211076)). For the South African case, this work’s results have found that if household work done predominantly by women were valued by its nearest specialized occupation, it would account for half of the nation’s aggregate labor income ([Oosthuizen, 2018](https://static1.squarespace.com/static/5994a30fe4fcb5d90b6fbeab/t/5baee9e8085229c69ad3840b/1538189818384/CWW+WP6.pdf)). Of course, if a given household activity can be equated to a specialized occupation, this implies that this activity comprises similar tasks to those performed in the specialized occupation. Then, by definition, the fact that tasks are an output of the application of a skill implies that at least some skills used outside of the market are comparable to those used in the market.
{% endhint %}

## Making Visible the "Unseen Economy"

For the “unseen economy,” i.e. activities outside of the specific SNA production boundary, we rely on an existing classification of all time-uses and tasks one may gain human capital from - the  [International Classification of Activities for Time Use Statistics (ICATUS)](https://unstats.un.org/unsd/classifications/Family/Detail/2083). This standard lists all time uses someone may have throughout the day. ICATUS represents the internationally applicable classifications of activities that people engage in during their 24-hour days.&#x20;

ICATUS is made up of three levels. The first digit level of disaggregation (highest level of aggregation) is called the major division, the second digit level is called the division level, and the third digit level, which is the most granular level, is called the group level. We first start by comparing ICATUS with the System of National Account, to define the boundaries of the "seen" and the "unseen" economies.&#x20;

<figure><img src="/files/eBF0pEEi3z8bEKnlpMG0" alt=""><figcaption><p>Mapping of seen and unseen economy across the SNA production boundary and ICATUS categories, general and context-specific application used by the 2011 South African Time Use Survey.  </p></figcaption></figure>

From the mapping above, it is evident that countries use the ICATUS taxonomy as a skeleton that they use to build a locally-adapted time use survey based on. Hence, a key characteristic that positions the ICATUS Framework very well for our purposes is its amenability to local contexts’ time use surveys.&#x20;

Based on the System of National Accounts, we therefore nominate the following ICATUS major division to encompass the unseen economy:&#x20;

| ICATUS Division 3                                                                          | ICATUS Division 4                                                           | ICATUS Division 5                                                                                                                                                         |
| ------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| *Unpaid domestic services for household and family members*                                | *Unpaid caregiving services for household and family members*               | *Unpaid volunteer, trainee and other unpaid work*                                                                                                                         |
| <mark style="color:red;">Excluding: 53. Unpaid trainee work and related activities.</mark> | <mark style="color:red;">Excluding: 9. Other unpaid work activities;</mark> | <mark style="color:red;">Excluding: All activities within ICATUS categories 3, 4 and 5 that include time allocated to waiting, traveling, or accompanying someone.</mark> |

## Matching "Unseen" ICATUS activities to the ESCO framework&#x20;

Once ICATUS activities have been associated with the "unseen" part of the economy, the challenge is to make sure the unseen part of our inclusive taxonomy follow the same structure as the existing ESCO taxonomy. Namely, we need to assign skills to the ICATUS activities forming the unseen part of the economy. Therefore, Tabiya conceptualizes a framework that links the most granular level of unseen economy ICATUS activities (3-digit level) to a set of non-exhaustive candidate ESCO skills and knowledge tags per ICATUS Activity.&#x20;

**ICATUS structure:** ICATUS 3 digit level activities comprise more specific activities that lie within ICATUS Divisions 3,4, and 5. For example, “Preparing meals/snacks'' is a Group Level Activity that falls within the Division 3 of ICATUS, namely Unpaid domestic services for household and family members. At the Group Level, ICATUS includes a definition of each activity, and a non-exhaustive list of the tasks that each activity includes and does not include, and at least one example of each Group Level Activity.&#x20;

**ESCO structure:** The ESCO taxonomy is made up of a set of occupations at level five or lower that are derived from the four ISCO-08 occupations hierarchy. For the purpose of the Tabiya Framework, we use only ESCO occupations at level five or lower for our analysis, since this provides the desirable granularity and metadata to adequately link ESCO to ICATUS. The ESCO taxonomy provides a brief description of each occupation it comprises, the occupation’s alternative labels as well as access to regulatory information pertaining to the occupation.&#x20;

{% hint style="success" %}
The ESCO taxonomy is particularly valuable because of how it structures the full set of skills required by an occupation: every occupation consists of "skills/competencies" and "knowledges" required, and each of these skills/competencies/knowledges are further assigned to being essential or optional to the occupation.&#x20;
{% endhint %}

<figure><img src="/files/8rrDNQPnIsCpGT1IucEc" alt=""><figcaption></figcaption></figure>

### **Mapping ICATUS x ESCO**

ICATUS Group Level activities are used as one of the inputs  and ESCO occupations and skills, competences and knowledge (skills, competences and knowledge will be referred to as skills henceforth for simplicity) are used as the other input for the matching procedure.&#x20;

First, each ICATUS Group Level activity is manually matched to at most four ESCO occupations by a team of researchers. ICATUS, even at the Group Level, is more broad than ESCO occupations and so, to ensure comprehensiveness in the Tabiya matching procedure, up to four ESCO occupations were potentially matched to each ICATUS Group Level activity. Since the ESCO taxonomy assigns skills to occupations, these ESCO-assigned skills then form the initial list of candidate assigned skills per ICATUS Group Level activity.&#x20;

{% hint style="info" %}
The assignment of ESCO occupations to ICATUS activities is based on the relative similarity definitions between ICATUS Group Level activities, and ESCO occupations. Comparing definitions of ICATUS activities to ESCO occupations is relatively easy to do amongst a team of researchers, since activities and occupations are similar in nature, and in turn, in the manner that each of these are defined.&#x20;
{% endhint %}

Once at most four ESCO occupations are matched to each ICATUS Group Level activity by a team of at least two researchers, these matches are then reviewed by another independent researcher. Throughout this process, there were no instances where the independent researcher disputed an occupation without mutual agreement from the pair of researchers who initially made the match. The process ensures that person-specific biases do not prevail during the matching process.&#x20;

{% hint style="danger" %}
While we endeavor to minimize biases and human error through this process, we cannot completely rule out the possibility that other (group level) biases and errors may have occured. Over time, as Tabiya accumulates more data and learns from this, the framework will use these learnings to improve and ultimately, not succumb to any currently prevailing biases and errors.
{% endhint %}

After the ICATUS to ESCO occupations match is done and verified, the next stage of the Tabiya framework formulation can commence. By design, the ESCO taxonomy pre-assigns skills to each ESCO occupation at disaggregation level five and below. Thus, we exploit this structure for the Tabiya framework. In particular, once ICATUS activities are matched to ESCO Occupations, we adopt the pre-assigned ESCO skills as the first set of candidate skills assigned to each unseen economy ICATUS activity. After duplicate skills are removed, each ICATUS activity now has a comprehensive list of skills assigned to it.  The procedure is shown below:

<img src="/files/no9f2X33j7E1VYpavTHm" alt="Matching process from ICATUS activities to ESCO skills" class="gitbook-drawing">

### Reducing the List of Skills

The skills taxonomy obtained following this method is hardly usable: each ICATUS activity ended up being associated with numerous skills. In order to provide a manageable list to job seekers when they are selecting skills, the team concluded that the list of skills needed to be condensed. Reducing this also allowed to addressed the transferability and signaling issues inherent to the unseen economy.&#x20;

We do this in two stages for our first use-case in the South African labor market: first, the Tabiya team proceeded with a first round of skill selection in order to delete skills that were deemed obviously not consistent with the definition and description of ICATUS activities. For the skills list that come out of this first stage filtering, a panel validation approach involving South-African labour market experts was adopted. An expert panel structured as a series of validation exercises would generate context-specific insights from a wide range of stakeholders, making it possible to compensate holders' biases pertaining to the transferability and credibility of skills.&#x20;

Below we describe the 4 principles used in the first stage skills reduction by Tabiya's researchers. The panel at Harambee is described in details [here](/tabiya-south-africa/unseen-economy/harambee-panel). &#x20;

1. **Isolating knowledges and retaining only skills/competencies**: The ESCO classification distinguishes between "skills" - that all describe an action - and "knowledges". For instance, the occupation "cook" is associated with the skill "use cooking techniques" and the knowledge "cooking technique". As a first step, we chose to isolate knowledges and to mainly focus on skills. However, we chose to let the panel decide to keep certain knowledges when they brought important new informations. For instance, "common children's diseases" may be deemed an essential knowledge when one claims to have experience raising children.&#x20;

{% hint style="info" %}
The decision to isolate knowledges also relies on the observation that knowledges assigned with ESCO occupations are usually redundant with skills, i.e. their presence in the list does not bring new information content. For instance, if a young job seeker "uses cooking techniques", this implies that they to know "cooking techniques".&#x20;
{% endhint %}

2. **Deleting irrelevant skills:**

   In ICATUS, each activity is associated with a definition, and explicit list of tasks included in the activity, a list of tasks that are not included, and one or more examples. ESCO is built similarly. However, ICATUS activities are typically not only broader, but conceptually different from ESCO occupations. For example, "Budgeting, planning, organizing duties and activities in the household" is not a formal sector activity that is included in ESCO. Therefore, to assign ESCO skills to this ICATUS activity, it was matched to "Office clerk", "Accountant", "Bed and breakfast operator", three ESCO occupations that encompass all tasks involved in "Budgeting, planning, organizing duties and activities in the household". \
   The issue is that they encompass a lot more work than what budgeting and planning for a household would entail. For example, because of "bed and breakfast operator", the skill "serve beverages" ended up being associated to  "Budgeting, planning, organizing duties and activities in the household", even though serving tasks are not included in the description of the ICATUS activity. When we came across such cases, we decided to exclude these skills. &#x20;

{% hint style="info" %}
The difficulty when applying this rule came from our expectation of how young job seekers would use the Harambee platform. Indeed, when selecting the ICATUS activity "preparing meals and snacks", one might mean that they prepare meals, serve them, and clean after. In ICATUS, those are three different activities, that are associated with different tasks, and thus skills. For version 0.1 of the Harambee platform, we chose to strictly fit the descriptions of each ICATUS activity, and to expect users to chose all relevant ICATUS activities. This choice was motivated by the observation that a more flexible approach would make a taxonomy inoperative, and make it more difficult to match job seekers to relevant job offers.&#x20;
{% endhint %}

3. **Deleting skills that are deemed too formal:** Some of the skills associated with ESCO occupations directly refer to situations or tasks only imaginable in the seen economy, whether it is formal or informal. For instance, "maintain customer service" cannot be appropriately associated with "serving meals and snacks", as it describes serving meals and snacks to one's own children/family members.&#x20;

{% hint style="info" %}
Applying this rule is entails having a very literal interpretation of ICATUS activities and the skills involved. For instance, one may argue that ensuring the satisfaction of family members when serving a meal or a snack may allow someone to develop customer service skills.&#x20;
{% endhint %}

4. **Deleting redundant skills:** ESCO occupations are typically associated with numerous skills, and each ICATUS activity was associated with multiple ESCO occupations. The lists of skills from each ESCO activity was added to a list of skills for each ICATUS activity, with deletion of duplicate instances of skills. For instance, "Outdoor cleaning" was associated with both "prune plants" and "prune hedges and trees". We considered that associating both skills to "outdoor cleaning" did not bring new information.&#x20;

{% hint style="info" %}
Applying this rule proved tricky, as it highlighted the complexity of the ESCO taxonomy. For instance, "prune plants" does not contain the subsoil "prune hedges and trees", even though hedges and trees are obviously plants. However, it contains the subskill "perform hand pruning". This shows that the organization of ESCO itself is not straightforward, and that cases exist where two skills conceptually very similar.&#x20;
{% endhint %}

This makes up the generalizable unseen economy framework methodology. From this candidate list of skills associated with each unseen economy ICATUS activity, differing contexts can begin to localize this framework.


# Why ESCO?

The European Skills, Competences, Qualifications, and Occupations (ESCO) taxonomy provides a comprehensive - albeit improvable - taxonomy of skills and occupations, on which our work builds.

## Our Challenge: Picking the Correct Base Taxonomy

As explained [here](/our-tech-stack/inclusive-livelihoods-taxonomy/methodology), Tabiya's work enhances widely used skill and competence taxonomies to ensure that our inclusive taxonomy aligns with existing tools employed by governments and organizations. Several key taxonomies serve as the foundation for job matching platforms and public policies, among which the ILO's International Standard Classification of Occupations (ISCO), the European Commission's European Skills, Competences, Qualifications and Occupations (ESCO; built on top of ISCO), and the United States' Occupation Information Network (O\*NET) are the most widely-accepted globally standardized taxonomies. Many countries have their own established taxonomies, built independently or based on these global frameworks such as South Africa's Organizing Framework for Occupations (OFO) and Singapore's SkillsFuture.&#x20;

Tabiya's challenge when picking a base taxonomy is the following: we were trying to pick a comprehensive base of the "seen" economy with which to be able to expand for the "unseen" economy.  This fundamental expansion of how global labor taxonomy is applied and can be used requires that the base taxonomy is -&#x20;

1. Broad enough to account for the occupations and human capital of individuals all over the world.&#x20;
2. Adaptable to local contexts, which implies using a language understandable by local users, and encompassing context-specific skills and occupations.&#x20;
3. Open-source and regularly updated, to account for the rapid evolution of labor markets, as well as the integration of new skills and occupations, such as the ones linked to AI or the green economy.&#x20;

Based on these considerations and discussions with our local implementing partners, ESCO was deemed the most appropriate taxonomy.&#x20;

## The Advantages of ESCO

### Accessibility to Job Seekers

1. When it comes to user experience, one of our main concerns is to make sure that demographics who may be traditionally marginalized from labor market intermediation solutions can use our inclusive taxonomy. In some contexts, the mastery of official languages, such as English in South Africa, is limited. Compared to O\*NET, ESCO relies on skill tags that are easier to understand.&#x20;
2. ESCO contains a list of ‘similar titles’ or what they call, "alternative labels" for each occupation – making it easier to search for an occupation from a broad starting point. For example, "data engineer" and "data expert" are saved as alternative labels for "data scientist".  &#x20;

### Highlighting the Market's and Job Seekers' (Updated) Skills

1. Contrary to O\*NET, ESCO includes "attitudes and values", which are soft-skills missing in O\*NET. Therefore, it covers a broader range of skills that are sought after by employers.&#x20;
2. ESCO is a European Commission project administered by the Directorate General Employment, Social Affairs and Inclusion. This ensures that the taxonomy is frequently updated and will be used by major actors in the long run. On the contrary, the last update of ISCO was made in 2008. Frequent updates of ESCO allow it to include skills and occupations associated with major evolutions in labor markets worldwide - for example, the ESCO taxonomy already contains designated skill trees for green and digital skills.  &#x20;

### Compatibility with Compass and Other Open-Source Tools

1. From a technical standpoint, ESCO is designed to be used by third-party organizations. It is specifically built for others to use and embed the taxonomy in their own technologies.&#x20;
2. ESCO is arguably more widely adopted across countries. Developing countries in [Latin America have also adopted ESCO as their skills framework rather than ONET](https://publications.iadb.org/publications/english/document/A-Skills-Taxonomy-for-LAC-Lessons-Learned-and-a-Roadmap-for-Future-Users.pdf), and they are contributing to the open-source tools around ESCO which can be leveraged (for instance,  ML-classification models which take in free-text job-titles and assign the ESCO occupations).
3. [ESCO provides a Local API in addition to their web-based API](https://ec.europa.eu/esco/portal/api). This means that Tabiya can host this API locally and adjust it as need be without depending on third-party performance when the volumes of requests are scaled up. This also gives Tabiya the ability to edit and adapt the framework freely.


# Lessons Learned

Extending the ESCO taxonomy by including the unseen economy while adapting our methodology to the challenges met gave us valuable lessons.

## ESCO vs O\*Net

{% hint style="info" %}
Under revision
{% endhint %}

## ICATUS Level Choice

As discussed [here](/our-tech-stack/inclusive-livelihoods-taxonomy), the ICATUS taxonomy comprises three levels that represent different levels of disaggregation, namely the major division, division and group levels. For Tabiya’s matching procedure, the group level activities pertaining to the unseen economy are matched to ESCO Occupations. There are a total of 60 ICATUS group level activities that comprise the Framework’s ICATUS input. The group level is chosen, rather than the major division or division level primarily due to the fact that the group level attains the highest level of precision in the ICATUS taxonomy. Thus, we are able to glean the most detail when matching from this level, relative to the major division and division levels.&#x20;

The challenge, though, has been in the matching of group level activities to ESCO occupations that occasionally define activities and occupations at differing levels of granularity. In particular, we find that there are some instances where an ESCO Occupation comprises more than one ICATUS group level. For instance, the ICATUS group level activities, “Cleaning up after food preparation/meals/snacks'' and “Preparing meals/snacks” are each matched to the ESCO occupation, “kitchen assistant” (defined as “Kitchen assistants assist in the preparation of food and cleaning of the kitchen area.”). Clearly, the occupation “kitchen assistant” is made up of these two ICATUS group level activities. Theoretically, one way to circumvent this would be to match ICATUS activities at the division level, rather than at the group level. However, this method is ultimately deemed too information costly since at the division level, the chance of omitting relatively rare but relevant ESCO occupations (and in turn, skills) from the ICATUS to ESCO match is very high. This does harm to the end user because it limits the eventual list of skills available to represent the work that they do. After discussion between the research team, and the Harambee team, we mutually agreed that the best way to do no harm in this case is to match ICATUS activities to ESCO occupations at the group level.

## ESCO occupations as a sub-step to skills (NLP, Word2Vec etc.)

The reason that ICATUS activities are matched to ESCO occupations, and then to skills as opposed to a match from ICATUS activities immediately to skills requires a summary of the conceptual roadmap taken to ultimately come to this decision. Initially, since the goal is to assign skills to activities, the most intuitive course of action seemed to be to directly match ICATUS activities to ESCO skills. In order to do this, Natural Language Processing (NLP) techniques were adopted. In particular, to compare ICATUS activities to ESCO skills, a technique called Word2Vec was adopted. Word2Vec is an algorithm used in NLP that uses word embeddings to assign numerical values to a given word in relation to the surrounding context of this word which is informed by the relationship that this given word has with other words in a dictionary used to train the algorithm (Mikolov et al., 2013). After several attempts to match ESCO skills directly to ICATUS activities, NLP techniques failed due the incompatibility in nature between ICATUS/ESCO as inputs and the required input for the algorithm to function optimally. Since ICATUS comprised short phrases pertaining to activities rather than single words as the ESCO taxonomy does for skills, Word2Vec could not optimally reconcile these taxonomies to produce a meaningful skills match. From this arose a precision-reproducibility tradeoff. On the one hand, NLP methods allow for near perfect reproducibility but score low on the precision scale in this context. A decision was taken to prioritize precision throughout the conceptualizing of the Tabiya Framework. Thus, the prevailing methodology that uses ESCO occupations as a crosswalk to skills proved to strike the precision-reproducibility scale optimally for our purposes. That is, the manual assignment of ESCO occupations to ICATUS activities to derive skills associated with unseen economy proved to score high on the precision scale and high enough on reproducibility scale, since this work was independently checked. Hence, the manual matching procedure prevailed for the Tabiya Framework’s creation.&#x20;

\
\ <br>


# Core Taxonomy

Our Core Taxonomy introduces important changes to ESCO to better reflect the livelihoods and economic realities in low- and middle-income countries, but avoids country-specific adaptations.

Our [Inclusive Livelihoods Taxonomy](/our-tech-stack/inclusive-livelihoods-taxonomy) introduces important changes to [ESCO](https://esco.ec.europa.eu/en) to better reflect the livelihoods and economic realities in low- and middle-income countries. This page presents our **standardized Core version, which is not country-specific.** We use our [Taxonomy Management Platform](/our-tech-stack/inclusive-livelihoods-taxonomy/open-taxonomy-platform)  to manage and maintain various adapted and country-specific versions.

{% hint style="info" %}
The Core version of our taxonomy is continuously evolving. Regular updates integrate feedback, corrections, and new insights from our partners to better represent diverse economic activities.
{% endhint %}

## Version 1.0.0

Version 1.0.0 of the Core Taxonomy expands conventional job and skill classifications by changing [ESCO v.1.1.1](https://esco.ec.europa.eu/en/about-esco/escopedia/escopedia/esco-v111) in the following ways:

* **Including New Occupations:** Such as "Micro-entrepreneur," explicitly recognizing self-employed individuals and small-scale entrepreneurial activities.
* **Capturing Unseen Economic Activities:** Integrating overlooked roles and activities such as caregiving, housework, volunteer work, and informal employment, guided by the [International Classification of Activities for Time-Use Statistics (ICATUS) 2016](https://unstats.un.org/unsd/classifications/Family/Detail/2083).
* **Enhancing Accessibility:** Adding simplified alternative labels and clearer terminology, making the taxonomy more approachable and easy to use.
* **Ensuring Data Quality:** Regularly refining taxonomy entries by fixing missing labels, removing duplicate records, and standardizing formatting and textual clarity.

### Detailed Release Notes

This version builds on [ESCO v1.1.1](https://esco.ec.europa.eu/en/about-esco/escopedia/escopedia/esco-v111). The latest Tabiya Core taxonomy version includes these specific updates:

* **Added:** New Micro-entrepreneur occupation (code: 5221\_2 uuid: d76f0bac-7145-4acc-a53e-259b96a6346b).
* **Added:** Unseen economy occupations/activities from the ICATUS 2016 classification:
  * 12 Occupation Groups (I3, I31, I32, I33, I34, I35, I36, I37, I4, I41, I42, I43, I44, I5, I51, I52).
  * 60 Occupations (e.g., I31\_0..., I32\_0..., I33\_0..., etc.).
  * 72 Occupation hierarchy entries.
  * 3689 Skills Relations.
* **Fixed:** Missing preferred labels in altLabels of ISCO groups, skill groups, some occupations, and some skills.
* **Fixed:** Duplicate alternative labels.
* **Fixed:** Invisible characters such as \t, nbsp, etc., in several text fields.
* **Fixed:** Leading/trailing spaces or new lines in several text fields.
* **Fixed:** Reformatted scope notes for better readability.

## File Format and Download

### CSV File Format — Compatible with ESCO

Our taxonomy is distributed in a CSV (Comma-Separated Values) format, ensuring ease of use, integration, and compatibility with existing ESCO-based systems.&#x20;

[A Zip file with all CSV files of Version 1.0.0 can be downloaded here.](https://platform.tabiya.tech/downloads/673b3fc9b52651611ad13a80-export-673b41d2b52651611ad13be4.zip)

#### Structure of the CSV Format:

* **`conceptId`:** A unique identifier (UUID) for each occupation or skill.
* **`preferredLabel`:** The primary, official label for the occupation or skill.
* **`alternativeLabels`:** Additional labels or synonyms, improving accessibility and understanding.
* **`description`:** Concise notes providing context, scope, or detailed explanations of occupations or skills.
* **`broaderConcept`:** Links to parent classifications, creating hierarchical relationships.
* **`narrowerConcept`:** Links to subordinate classifications, detailing hierarchical depth.
* **`relatedConcept`:** Associations indicating related occupations or skills.
* **`skillRelations`:** Explicit relationships linking specific skills to relevant occupations, allowing precise mapping.
* **`iscoGroup`:** ISCO occupational group codes aligning occupations with international standards.

This format closely follows ESCO’s CSV structure, enhancing interoperability and facilitating easy integration into existing workflows. For detailed technical documentation, refer to the [Tabiya CSV Format Specification](#csv-file-format-compatible-with-esco).

## Licensing and Availability

The Tabiya Core Taxonomy is openly available under the **Creative Commons Attribution 4.0 International (CC BY 4.0)** license, facilitating unrestricted use, adaptation, and sharing with appropriate attribution.


# Open Taxonomy Platform

Labor markets are diverse and fast changing, so a useful taxonomy must be localizable and adaptable in a transparent way. Our open taxonomy platform empowers partners to do this.

We are currently developing the **Open Taxonomy Platform,** which lets partners flexibly adapt our reference taxonomy in a transparent way. The platform is currently under active open-source development. You can [follow along and contribute on Github](https://github.com/tabiya-tech/taxonomy-model-application). If you would like to learn more, please [get in touch](mailto:hi@tabiya.tech).

## Key Features

1. **Reference Taxonomy**\
   We maintain a canonical set of occupations and skills for general use. This includes both traditional roles (e.g., *Accountant*) and informal or unpaid roles (e.g., *Family Caregiver*).
2. **Localization**\
   Partners create “forks” or branches of the base taxonomy to adapt local job titles, language requirements, or cultural specifics—without losing the broader structure. This means a localized taxonomy can still talk “under the hood” to other regions.
3. **API Access**\
   A simple REST API (with plans for GraphQL in the future) allows you to search, retrieve, and update taxonomy entries. Whether you’re building a job-matching platform or a chatbot, you can programmatically look up occupations, skills, tasks, or synonyms. An overview of the open taxonomy platform API is available [here](/our-tech-stack/inclusive-livelihoods-taxonomy/open-taxonomy-platform/open-taxonomy-platform-api).
4. **Version Control & Transparency**\
   Every revision is tracked, and older versions remain accessible. Anyone can see what changed, when, and why—a key requirement when multiple organizations or contributors collaborate over time.
5. **Community Contributions**\
   We welcome pull requests and issue reports on GitHub. Contributors can:
   * Propose new occupations or skills that capture informal work or emerging digital livelihoods.
   * Offer local language translations.
   * Flag redundancies or gaps in the taxonomy structure.

### Use Cases

This taxonomy is best used for:&#x20;

* **Public Employment Services (PES)**\
  A government job portal may need to index both formal occupations (like *Nurse* or *Electrician*) and informal tasks (like *Childcare at Home*). The Open Taxonomy Platform lets PES administrators adapt the reference taxonomy to local categories while remaining compatible with international frameworks.
* **Youth Employment NGOs**\
  A non-profit that focuses on upskilling young people can integrate the platform’s API into its app. When a user describes an unpaid or informal experience—such as cooking at home for large extended families—the NGO’s app can quickly map that to recognized culinary or budget-management skills.
* **Research and Analytics**\
  Think tanks or academic teams might want to track how the labor market shifts over time, especially in informal sectors. By using version-controlled, transparent data on occupations and skills, they can compare how certain roles or skill sets grow or shrink across multiple regions.

### Getting Started

1. **Explore the Reference Taxonomy**\
   Head to our [GitHub repository](https://github.com/tabiya-tech/taxonomy-model-app) for an overview. You’ll find JSON definitions of occupations, tasks, and skill mappings, along with instructions on how to propose changes.
2. **Set Up the Platform Locally**\
   If you want to adapt or extend the taxonomy for your own region, clone the repo and follow the quick-start guide to run a local instance of our platform.
3. **Use the API**\
   An overview of the Open Taxonomy Platform API is available [here](/our-tech-stack/inclusive-livelihoods-taxonomy/open-taxonomy-platform/open-taxonomy-platform-api). Further documentation on authentication, queries, and updates is available in our `/docs` folder on GitHub. You can also find sample scripts and postman collections that show how to retrieve occupations, look up skill definitions, and commit changes.
4. **Contribute**
   * **Issues and Pull Requests**: Found an occupation that needs local adaptation or noticed a skill that’s missing? Open an issue or submit a pull request on GitHub.
   * **Discussion & Feedback**: Join our community forum (link coming soon!) where practitioners and developers share best practices, tips, and localized expansions.

### Roadmap

While the Open Taxonomy Platform is already in active development, some exciting features are in progress:

* **Multilingual Support**: We aim to standardize how we store localized names and synonyms for occupations and skills, making it easy to deploy in multiple languages.
* **Advanced API Queries**: Searching by skill clusters or skill synonyms to quickly find relevant occupations.
* **Integrations with Compass and Classifier**: Our [Compass conversational tool](https://docs.tabiya.org/compass) and Inclusive Livelihoods Classifier both rely on the taxonomy. We plan to release more examples of how to combine them seamlessly.

### License and Community Guidelines

The Open Taxonomy Platform is licensed under MIT, reflecting our commitment to openness and widespread adoption. We encourage everyone—government agencies, NGOs, private job-matching platforms, researchers, and individuals—to explore, test, and refine the taxonomy so that, together, we build a more inclusive, dynamic picture of labor markets.


# Open Taxonomy Platform API

Labor markets are diverse and fast changing, so a useful taxonomy must be localizable and adaptable in a transparent way. Learn more about the API that powers our open taxonomy platform.

## Overview and Format

The [Open Taxonomy Platform](/our-tech-stack/inclusive-livelihoods-taxonomy/open-taxonomy-platform) API provides secure access to taxonomy models, occupations, skills, and their respective groups, enabling seamless integration with your applications. The base URL for all API requests is [https://taxonomy.tabiya.tech/taxonomy/api-doc/swagger/](https://taxonomy.tabiya.tech/api-doc/swagger/). To ensure standardized communication, all requests and responses (except file uploads) utilize the JSON format, making it straightforward to integrate with any modern programming environment.

## Credentials and Authentication

#### Prerequisites

Before you can integrate with the APIs, you must obtain credentials from the platform administrators. **Request credentials** via the dedicated email address on this [page](https://docs.tabiya.org/#discover-tabiyas-work). Depending on the authentication method you choose, you will receive:&#x20;

* **API Keys**: A unique `X-API-Key`&#x20;
* **M2M OAuth:** An `Authorization server URL`, `Client ID`, and `Client Secret`

#### API Path Prefixes

All partner APIs use the following path prefix:

* `/api/partner`  for API Keys
* `/api/app`  For JWT Tokens received via M2M OAuth.

#### Authentication Methods

We support two authentication methods. Choose the one that aligns with your security requirements.

#### API Keys

API Keys provide a simple authentication mechanism suitable for basic integrations.

**Usage**

Include the API key in every request using the `X-API-Key` HTTP header.

**Example:-**

```bash
curl -X GET \
  https://taxonomy.tabiya.tech/api/partner/info \
  -H "X-API-Key: YOUR_API_KEY"
```

#### Machine to Machine (M2M) OAuth 2.0

Machine to Machine OAuth 2.0 is the recommended method for secure, automated service-to-service communication using short-lived access tokens

**Step 1**

Send an HTTP `POST` request to the **Authorization server URL** using the `Client ID` and `Client Secret`.

**Example**

```bash
curl -X POST YOUR_AUTHORIZATION_SERVER_URL \
     -H "Content-Type: application/x-www-form-urlencoded" \
     -d "grant_type=client_credentials&client_id=YOUR_CLIENT_ID&client_secret=YOUR_CLIENT_SECRET"
```

The authorization server responds with an access token:

```json
{
   "access_token": "YOUR_ACCESS_TOKEN",
   "token_type": "Bearer",
   "expires_in": 3600
}
```

For more information refer to the Auth token documentation on the part of exchanging client credentials for [access token](https://oauth.net/2/access-tokens/).

* <https://docs.aws.amazon.com/cognito/latest/developerguide/token-endpoint.html#post-token-positive-exchanging-client-credentials-for-an-access-token-in-request-body>

**Step 2**

Send a request to the API with the access token from the previous result.

```bash
curl -X GET \
   https://taxonomy.tabiya.tech/api/app/info \
   -H "Authorization: Bearer YOUR_ACCESS_TOKEN"
```

For more details, see the [Scopes, M2M, and resource servers](https://docs.aws.amazon.com/cognito/latest/developerguide/cognito-user-pools-define-resource-servers.html) documentation on AWS.

## Direct Access to the OpenAPI Specification

Two access methods are available for the API specifications. The primary method is through the interactive Swagger UI documentation, which allows browsing and live testing of all endpoints. Crucially, for developers looking to quickly configure their clients, a direct, standalone link to the OpenAPI v3 JSON specification file is provided. This link is essential for developers who want to import the entire API definition into tools like Postman or Insomnia (using the *Import > Link option*), accelerating setup by automatically generating all endpoint, parameter, and schema definitions:

* OpenAPI Specification link: <https://taxonomy.tabiya.tech/api-doc/swagger/tabiya-api.json>
* Swagger UI link: <https://taxonomy.tabiya.tech/api-doc/swagger/>
* ReDoc UI link: <https://taxonomy.tabiya.tech/api-doc/redoc/>


# Taxonomy CSV Format

The Tabiya CSV format is used to import and export data from the Tabiya Open Taxonomy platform.

Each taxonomy version is made up of nine CSV files Each file contains a different type of data. The files are:

* [Model Info](#model-info)
* [Skill Groups](#skill-groups)
* [Skills](#skills)
* [Skill Hierarchy](#skill-hierarchy)
* [Skill to Skill Relations](#skill-to-skill-relations)
* [Occupation Groups](#occupation-groups)
* [Occupations](#occupations)
* [Occupation Hierarchy](#occupation-hierarchy)
* [Occupation to Skill Relations](#occupation-to-skill-relations)
* [LICENSE](#LICENSE)

## General notes on the fields of the CSV files

### UUID History

A `UUIDHISTORY` field is a [list](#lists) of all the UUIDs that have been assigned to an entity during its lifecycle, e.g. when the entity is created, imported, exported or copied into our platform.

It is an identifier that can be used for tracking objects not only across their lifecycle, but also across systems.

The UUID history is ordered from newest to oldest UUID.

The first entry in the list is the current UUID of the object.\
The last entry in the list is the very first (*initial*) UUID of the object.

The entities in this dataset have been assigned an initial UUID. When an entity is imported into our platform, a new UUID will be issued and added at the top of UUID history.

The UUID used by the platform are based on the [Universally Unique Identifier v4](https://datatracker.ietf.org/doc/html/rfc4122) standard.

> The maximum number of UUIDs in the history for an object is constrained to `10000`.

### Origin Uri

The `ORIGINURI` field is a [URI](https://datatracker.ietf.org/doc/html/rfc3986) that points to the location where an entity was originally defined.

> The maximum length for the Origin Uri is `4096` characters.

### ID

The `ID` field is a unique identifier for each entity in the CSV dataset. It is used for referencing within the CSV dataset, for example, in the relations between entities.

This field is not meant to be used as an identifier outside the scope of the CSV files, for that purpose you should use the first entry in the [UUID History](#uuid-history).

### Object Types

The object types are used to differentiate between different types of entities in the dataset.

For example in relations between entities, the object types are used to specify the type of the parent and child objects and determine in which file these objects can be located.

The object types in the CSV files are:

* `skill`: Represents a [skill](#skills).
* `skillgroup`: Represents a [skill group](#skill-groups).
* `escooccupation`: Represents an [occupation](#occupations) that originates from the ESCO framework.
* `localoccupation`: Represents an [occupation](#occupations) that not originate from the ESCO framework and is defined only this taxonomy.
* `occupationgroup`: Represents an [Occupation group](#occupation-groups).

### Lists

List properties are stored in the CSV files as strings separated by a  character. Currently, we do not support values that contain a new line.

### Dates

The dates in the CSV files are stored in the [ISO 8601](https://www.iso.org/iso-8601-date-and-time-format.html) format.

## File descriptions

### Model Info

Contains information about the model. The export filename is `model_info.csv`

#### Columns

* [`UUIDHISTORY`](#uuid-history): A list of [UUIDs](#uuid-history).
* `NAME`: The name of the model.
* `LOCALE`: The short code of the model's locale.
* `DESCRIPTION`: The description of the model.
* `VERSION`: The version of the model.
* `RELEASED`: A boolean value that indicates whether the model is released or not.
* `RELEASENOTES`: The release notes of the model.
* `CREATEDAT`: The [date](#dates) the model was created.
* `UPDATEDAT`: The [date](#dates) the model was last updated.

### Skills

Contains the skills of the taxonomy. The export filename is `skills.csv`

#### Columns

* [`ORIGINURI`](#origin-uri): A [URI](#origin-uri) that points to the location where the skill was originally defined.
* [`ID`](#id): A [unique identifier](#id), used for referencing the skill within the CSV dataset.
* [`UUIDHISTORY`](#uuid-history): A list of [UUIDs](#uuid-history).
* `SKILLTYPE`: The skill type.
  * Possible values: `skill/competence`,`knowledge`,`language`,`attitude` or empty ( ).
* `REUSELEVEL`: The skill reuse level.
  * Possible values: `sector-specific`,`occupation-specific`,`cross-sector`,`transversal` or empty ( ).
* `PREFERREDLABEL`: The preferred label of the skill.
* `ALTLABELS`: A [list](#lists) of alternative labels for the skill.
  * Maximum length per label: `256` characters.
  * Maximum number of labels: `100`.
* `DESCRIPTION`: The skill description.
  * Maximum length:`4000` characters.
* `DEFINITION`: The skill definition.
  * Maximum length:`4000` characters.
* `SCOPENOTE`: The skill scope note.
  * Maximum length:`4000` characters.
* `ISLOCALIZED`: A boolean value that indicates whether the skill is localized or not.
  * Possible values: `true` or `false`.
* `CREATEDAT`: The [date](#dates) the skill was created.
* `UPDATEDAT`: The [date](#dates) the skill was last updated.\\

### Skill Groups

Contains the skill groups of the taxonomy. The export filename is `skill_groups.csv`

#### Columns

* [`ORIGINURI`](#origin-uri): A [URI](#origin-uri) that points to the location where the skill group was originally defined.
* [`ID`](#id): A [unique identifier](#id), used for referencing the skill group within the CSV dataset.
* [`UUIDHISTORY`](#uuid-history): A list of [UUIDs](#uuid-history).
* `CODE`: SkillGroup code as defined in ESCO. It has the general format `SX.X.X`, where `X` is a number.
* `PREFERREDLABEL`: The preferred label of the skill group.
* `ALTLABELS`: A [list](#lists) of alternative labels for the skill group.
  * Maximum length per label: `256` characters.
  * Maximum number of labels: `100`.
* `DESCRIPTION`: The skill group description.
  * Maximum length:`4000` characters.
* `SCOPENOTE`: The skill group scope note.
  * Maximum length:`4000` characters.
* `CREATEDAT`: The [date](#dates) the skill group was created.
* `UPDATEDAT`: The [date](#dates) the skill group was last updated.

### Occupations

Contains the occupations of the taxonomy. The export filename is `occupations.csv`

#### Columns

* [`ORIGINURI`](#origin-uri): A [URI](#origin-uri) that points to the location where the occupation was originally defined.
* [`ID`](#id): A [unique identifier](#id), used for referencing the occupation within the CSV dataset.
* [`UUIDHISTORY`](#uuid-history): A list of [UUIDs](#uuid-history).
* `OCCUPATIONGROUPCODE`:The Occupation group that the occupation belongs to.
* `CODE`: An occupation code assigned to the occupation.
  * For ESCO occupations, the code will be the parent code, followed by a `.` and any number of digits. Eg: `XXXX.1234`
  * For local occupations, the code will be the parent code, followed by an `_` and any number of digits. `XXXX_1234`
* `PREFERREDLABEL`: The preferred label of the occupation.
* `ALTLABELS`: A [list](#lists) of alternative labels for the occupation.
  * Maximum length per label: `256` characters.
  * Maximum number of labels: `100`.
* `DESCRIPTION`: The occupation description.
  * Maximum length:`4000` characters.
* `DEFINITION`: The occupation definition.
  * Maximum length:`4000` characters.
* `SCOPENOTE`: The occupation scope note.
  * Maximum length:`4000` characters.
* `REGULATEDPROFESSIONNOTE`: The regulated profession note.
  * Maximum length:`4000` characters.
* `OCCUPATIONTYPE`: The type of the occupation.
  * Possible values: `escooccupation` or `localoccupation`.
* `ISLOCALIZED`: A boolean value that indicates whether the occupation is localized or not. Only ocuppations of the type `escooccupation` can be localized.
  * Possible values: `true` or `false`.
* `CREATEDAT`: The [date](#dates) the occupation was created.
* `UPDATEDAT`: The [date](#dates) the occupation was last updated.

### Occupation Groups

Contains the Occupation groups of the taxonomy. The export filename is `occupation_groups.csv`

### Columns

* [`ORIGINURI`](#origin-uri): A [URI](#origin-uri) that points to the location where the Occupation group was originally defined.
* [`ID`](#id): A [unique identifier](#id), used for referencing the Occupation group within the CSV dataset.
* [`UUIDHISTORY`](#uuid-history): A list of [UUIDs](#uuid-history).
* `CODE`: A four digit identification code of the Occupation group. Each digit represents a level in the hierarchy.
  * For ISCO groups, the code is a maximum of 4 digits, and each child group should have a code that begins with the parent group code. Eg: `1234`
  * For local groups without a parent group, the code should start with an alphabetical character. Eg: `A1234`
  * For local groups, if the parent occupation group is an isco group, the code should start with the parent group code and then have one alphabetical character. Eg: `1234A`
  * For local groups, if the parent occupation group is also a local group, the code should start with the parent group code and then have either an alphabetical character or a number. Eg: `1234AB` or `1234A1`
* `GROUPTYPE`: The type of the Occupation group.
  * Possible values: `iscogroup` or `localgroup`.
* `PREFERREDLABEL`: The preferred label of the Occupation group.
* `ALTLABELS`: A [list](#lists) of alternative labels for the Occupation group.
  * Maximum length per label: `256` characters.
  * Maximum number of labels: `100`.
* `DESCRIPTION`: The Occupation group description.
  * Maximum length:`4000` characters.
* `CREATEDAT`: The [date](#dates) the Occupation group was created.
* `UPDATEDAT`: The [date](#dates) the Occupation group was last updated.

### Skill-to-Skill Relations

Contains the relations between skills. The export filename is `skill_to_skill_relations.csv`

#### Columns

* `REQUIRINGID`: The [`ID`](#id) of the skill that requires another skill.
* `RELATIONTYPE`: The type of the relation.
  * Possible values: `essential` or `optional`.
* `REQUIREDID`: The [`ID`](#id) of the skill that is required by another skill.
* `CREATEDAT`: The [date](#dates) the relation was created.
* `UPDATEDAT`: The [date](#dates) the relation was last updated.

### Occupation-to-Skill Relations

Contains the relations between occupations and skills. The export filename is `occupation_to_skill_relations.csv`

#### Columns

* `OCCUPATIONTYPE`: The type of the occupation.
  * Possible values: `escooccupation` or `localoccupation`.
* `OCCUPATIONID`: The [`ID`](#id) of the occupation.
* `RELATIONTYPE`: The type of the relation.
  * Possible values: `essential`, `optional`, or it can be left empty.
* `SIGNALLINGVALUELABEL`: The signalling value label of the relation.
  * Possible values: `low`, `medium`, `high`, or it can be left empty.
* `SIGNALLINGVALUE`: The signalling value of the relation.
  * A number between `0` and `1`, or it can be left empty. The only allowed delimiter for decimal numbers is a `.`.
* `SKILLID`: The [`ID`](#id) of the skill.
* `CREATEDAT`: The [date](#dates) the relation was created.
* `UPDATEDAT`: The [date](#dates) the relation was last updated.

> Caveat: An escooccuption cannot have a `signalling value` or `signalling value label`. It **must** have a `relationType`.\
> For localoccupations `signalling value` and `relationType` are mutually exclusive. A `localoccupation` can **either** have a `signalling value` and `signalling value label` **or** it can have a `relationType`, but not both.

### Skill Hierarchy

Contains the hierarchical structure of various skills. The export filename is `skill_hierarchy.csv`

#### Columns

* `PARENTOBJECTTYPE`: The type of the parent object.
  * Possible values: `skill` or `skillgroup`.
* `PARENTID`: The [`ID`](#id) of the parent object.
* `CHILDID`: The [`ID`](#id) of the child object.
* `CHILDOBJECTTYPE`: The type of the child object.
  * Possible values: `skill` or `skillgroup`.
* `CREATEDAT`: The [date](#dates) the relation was created.
* `UPDATEDAT`: The [date](#dates) the relation was last updated.

> Caveat: A skill cannot be the parent of a skill group.

### Occupation Hierarchy

Contains the hierarchical structure of various occupations. The export filename is `occupation_hierarchy.csv`

#### Columns

* `PARENTOBJECTTYPE`: The type of the parent object.
  * Possible values: `occupationgroup`, `escooccupation`, `localoccupation`.
* `PARENTID`: The [`ID`](#id) of the parent object.
* `CHILDID`: The [`ID`](#id) of the child object.
* `CHILDOBJECTTYPE`: The type of the child object.
  * Possible values: `occupationgroup`, `escooccupation`, `localoccupation`.
* `CREATEDAT`: The [date](#dates) the relation was created.
* `UPDATEDAT`: The [date](#dates) the relation was last updated.

> Caveat: An `escooccupation` cannot be the parent of an 'occupationgroup'.\
> Caveat: An `localoccupation` can be a child of an `escooccupation` or another `localoccupation`.

### LICENSE

Contains the license information for the model. If one wants to add a license to the dataset, it can be added to a file named `LICENSE` in the root of the dataset.\
The `LICENSE` file supports plain text and Markdown format. During export the license information of the model will also be exported in the `LICENSE` file.


# Livelihoods Classifier

The Tabiya Livelihoods Classifier provides an easy-to-use implementation of the entity-linking paradigm to support job description heuristics.

Using state-of-the-art transformer neural networks this tool can extract 5 entity types: Occupation, Skill, Qualification, Experience, and Domain. For the Occupations and Skills, ESCO-related entries are retrieved. The procedure consists of two discrete steps, entity extraction and similarity vector search.

### Model's Architecture

<figure><img src="/files/e8Yt8cexDoR01khu47Cr" alt=""><figcaption><p>Job Entity Linking Pipeline</p></figcaption></figure>

### Target Audience

The tool is intended for specialists in workforce analytics, recruitment technologies, and human capital management, as well as researchers focused on labor markets and job data. It caters to organizations and professionals who deal with analyzing, categorizing, and optimizing job descriptions or resumes at scale. Ideal users include those in HR technology, career advisory services, and data-driven talent solutions, particularly those seeking to enhance precision in identifying roles, skills, and qualifications for improved decision-making in hiring or workforce planning. Technical users interested in integrating standardized frameworks into their analysis will also find this tool highly relevant.

### Related research

Our Livelihoods Classifier's design and evaluation is presented in Saroglou, S., Diamantaras, K., Preta, F., Delianidi, M., Benisis, A. & Meyer, C.J. (2025). Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies. *arXiv preprint arXiv:2512.03195*. <https://arxiv.org/abs/2512.03195>.

{% hint style="info" %}
You can find all related code on [Tabiya's GitHub page](https://github.com/tabiya-tech/tabiya-livelihoods-classifier/tree/main).
{% endhint %}


# Getting Started

## Installation

Prerequisites\\

* A recent version of [git](https://git-scm.com/) (e.g. ^2.37 )
* [Python 3.10 or higher](https://www.python.org/downloads/)
* [Poerty 1.8 or higher](https://python-poetry.org/)

  > Note: to install Poetry consult the [Poetry documentation](https://python-poetry.org/docs/#installing-with-the-official-installer)

  > Note: Install poetry system-wide (not in a virtualenv).
* [Git LFS](https://git-lfs.github.com/)

#### Using Git LFS

This tool uses Git LFS for handling large files. Before using it you need to install and set up Git LFS on your local machine. See <https://git-lfs.com/> for installation instructions.

After Git LFS is set up, follow these steps to clone the repository:

```shell
git clone https://github.com/tabiya-tech/tabiya-livelihoods-classifier.git
```

If you already cloned the repository without Git LFS, run:

```shell
git lfs pull
```

#### Install the dependencies <a href="#dep" id="dep"></a>

**Set up virtualenv**

In the **root directory** of the backend project (so, the same directory as this README file), run the following commands:

```shell
# create a virtual environment
python3 -m venv venv

# activate the virtual environment
source venv/bin/activate
```

```shell
# Use the version of the dependencies specified in the lock file
poetry lock --no-update
# Install missing and remove unreferenced packages
poetry install --sync
```

> Note: Install the dependencies for the training using:
>
> ```shell
> # Use the version of the dependencies specified in the lock file
> poetry lock --no-update
> # Install missing and remove unreferenced packages
> poetry install --sync --with train
> ```

> Note: Before running any tasks, activate the virtual environment so that the installed dependencies are available:
>
> ```shell
> # activate the virtual environment
> source venv/bin/activate
> ```
>
> To deactivate the virtual environment, run:
>
> ```shell
> # deactivate the virtual environment
> deactivate
> ```

Activate Python and download the NLTK punctuation package to use the sentence tokenizer. You only need to download `punkt` it once.

```shell
python <<EOF
import nltk
nltk.download('punkt')
EOF
```

#### Environment Variable & Configuration

The tool uses the following environment variable:

* `HF_TOKEN`: To use the project, you need access to the HuggingFace 🤗 entity extraction model. Contact the administrators via \[<tabiya@benisis.de>]. From there, you must create a read access token to use the model. Find or create your read access token [here](https://huggingface.co/settings/tokens). The backend supports the use of a `.env` file to set the environment variable. Create a `.env` file in the root directory of the backend project and set the environment variables as follows:

```dotenv
# .env file
HF_TOKEN=<YOUR_HF_TOKEN>
```

> ATTENTION: The .env file should be kept secure and not shared with others as it contains sensitive information.

## QuickStart Guide

## Inference Pipeline

The inference pipeline extracts occupations and skills from a job description and matches them to the most similar entities in the ESCO taxonomy.

### Usage

First, activate the virtual environment as explained [here](#dep).

Then, `start python interpreter in the root directory` and run the following commands:

Load the `EntityLinker` class and create an instance of the class, then perform inference on any text with the following code:

```python
from inference.linker import EntityLinker
pipeline = EntityLinker(k=5)
text = 'We are looking for a Head Chef who can plan menus.'
extracted = pipeline(text)
print(extracted)
```

After running the commands above, you should see the following output:

```js
[
  {'type': 'Occupation', 'tokens': 'Head Chef', 'retrieved': ['head chef', 'industrial head chef', 'head pastry chef', 'chef', 'kitchen chef']},
  {'type': 'Skill', 'tokens': 'plan menus', 'retrieved': ['plan menus', 'plan patient menus', 'present menus', 'plan schedule', 'plan engineering activities']}
]
```

### French version

You can use the French version of the Entity Linker using the following code:

```python
from inference.linker import FrenchEntityLinker
pipeline = FrenchEntityLinker(entity_model = 'tabiya/camembert-large-job-ner', similarity_model = 'intfloat/multilingual-e5-base')

text = 'Nous recherchons un chef de cuisine capable de planifier les menus.'
extracted = pipeline(text)
print(extracted)
```

You should see the following output:

```javascript
[
  {'type': 'Occupation', 'tokens': 'chef de cuisine', 'retrieved': ['chef de cuisine', 'chef de marque', 'chef mécanicien', 'chef cuisinier/cheffe cuisinière', 'chef de train']}, 
  {'type': 'Skill', 'tokens': 'planifier les menus', 'retrieved': ['planifier les menus', 'présenter des menus', 'établir les menus des patients', 'préparer des plannings', 'préparer des plats préparés']}
]
```

### Running the evaluation tests

Load the `Evaluator` class and print the results:

```python
from inference.evaluator import Evaluator

results = Evaluator(entity_type='Skill', entity_model='tabiya/roberta-base-job-ner', similarity_model='all-MiniLM-L6-v2', crf=False, evaluation_mode=True)
print(results.output)
```

This class inherits from the `EntityLinker`, with the main difference being the `'entity_type'` flag.

{% hint style="warning" %}
If you want to run evaluations on custom datasets, you will need to make modifications to the `_load_dataset` function, located on the `evaluation.py` file. Please refer to the original evaluation datasets as described [here](/our-tech-stack/livelihoods-classifier/datasets). If you have any trouble, please open an issue on [GitHub](https://github.com/tabiya-tech/tabiya-livelihoods-classifier/issues).
{% endhint %}

### Minimum Hardware

* 4 GB CPU/GPU RAM

The code runs on GPU if available. Ensure your machine has CUDA installed if running on GPU.


# Web Application

For ease of use, we developed a simple FullStack application (a Flask-based API as a BackEnd and jQuery FrontEnd) to analyze job descriptions and predict relevant occupations, skills, and qualifications using the entity-linking model.

### Usage

First, activate the virtual environment as explained [here](/our-tech-stack/livelihoods-classifier/getting-started#dep). Then, run the following command in python in the `root` directory:

#### Running the API

**Run the Flask application**:

```bash
python app/server/matching.py
```

Or set the Flask application environment variable and use the Flask command:

```bash
export FLASK_APP=app/server/matching.py
flask run --host=0.0.0.0 --port=5001
```

### Example Usage

1. **Open the browser** and navigate to `http://127.0.0.1:5001/`.
2. **Paste a job description** into the provided text area.
3. **Click the "Analyze Job" button** to send the job description to the `/match` endpoint.
4. **View the results** under "Predicted Occupations," "Predicted Skills," and "Predicted Qualifications."

{% hint style="danger" %}
This app is just for demonstration purposes. If you wish to deploy this model, use a more reliable/secure strategy.
{% endhint %}


# Datasets

## Reference Sets

#### Occupations

* **Location**: inference/files/occupations\_augmented.csv
* **Source**: [ESCO dataset - v1.1.1](https://esco.ec.europa.eu/en/use-esco/download)
* **Description**: ESCO (European Skills, Competences, Qualifications and Occupations) is the European multilingual classification of Skills, Competences, and Occupations. This dataset includes information relevant to the occupations.
* **License**: Creative Commons Attribution 4.0 International see DATA\_LICENSE for details.
* **Modifications**: The columns retained are `alt_label`, `preferred_label`, `esco_code`, and `uuid`. Each alternative label has been separated into individual rows.

#### Skills

* **Location**: inference/files/skills.csv
* **Source**: [ESCO dataset - v1.1.1](https://esco.ec.europa.eu/en/use-esco/download)
* **Description**: ESCO (European Skills, Competences, Qualifications and Occupations) is the European multilingual classification of Skills, Competences and Occupations. This dataset includes information relevant to the skills.
* **License**: Creative Commons Attribution 4.0 International see Data License for details.
* **Modifications**: The columns retained are `preferred_label` and `uuid`.

#### Qualifications

* **Location**: inference/files/qualifications.csv
* **Source**: [Official European Union EQF comparison website](https://europass.europa.eu/en/compare-qualifications)
* **Description**: This dataset contains EQF (European Qualifications Framework) relevant information extracted from the official EQF comparison website. It includes data strings, country information, and EQF levels. Non-English text was ignored.
* **License**: Please refer to the original source for [license information](https://europass.europa.eu/en/node/2161).
* **Modifications**: Non-English text was removed, and the remaining information was formatted into a structured database.

For the French version of the tool, we use the French version of ESCO v1.1.1, as well as, a translation of the qualifications, using the Google Translation API.

## Training Sets

#### **Entity Extraction**

* **Location:** [job\_ner\_dataset](https://huggingface.co/datasets/tabiya/job_ner_dataset)
* **Source:** [Green Benchmark corpus](https://github.com/acp19tag/skill-extraction-dataset)
* **Description:** This dataset provides a comprehensive benchmark suite for Entity Recognition (ER) in job descriptions. Developed to fill the significant gap in resources for extracting key entities like skills from job descriptions, the dataset features 18.6k annotated entities across five categories: Skill, Qualification, Experience, Occupation, and Domain.
* **License:** CC-BY-NC-4.0
* **Modifications:** No modifications were made to the original dataset. It was only converted to HuggingFace format.

#### **Entity Similarity**

* **Location:** TBD
* **Source:**[ hahu-occupation-titles](https://huggingface.co/datasets/tabiya/occupation_titles_esco)
* **Description:**

  The `hahu_test.csv` file is the original file provided by Hahu Jobs with the following fields:

  * title: The title of the job position, indicating the specific role and/or position within the organization.
  * esco\_label: The preferred or alternative label provided by ESCO, matching the corresponding ESCO code.
  * esco\_code: The ESCO code associated with the job, facilitating standardized classification and comparison across different job listings.
* **License:** CC-BY-NC-4.0
* **Modifications:** Extracted Occupation title and relevant ESCO code and matched with preferred and alternative labels.

## Evaluation Sets

#### Hahu Test

* **Location**: inference/files/eval/redacted\_hahu\_test\_with\_id.csv
* **Source**: [hahu\_test](https://huggingface.co/datasets/tabiya/hahu_test)
* **Description**: This dataset consists of 542 entries chosen at random from the 11 general classification system of the Ethiopian hahu jobs platform. 50 entries were selected from each class to create the final dataset.
* **License**: Creative Commons Attribution 4.0 International see Data License for details.
* **Modifications**: No modifications were made to the selected entries.

#### House and Tech

* **Location**:
  * inference/files/eval/house\_test\_annotations.csv
  * inference/files/eval/house\_validation\_annotations.csv
  * inference/files/eval/tech\_test\_annotations.csv
  * inference/files/eval/tech\_validation\_annotations.csv
* **Source**: Provided by [Decorte et al.](https://arxiv.org/abs/2209.05987)
* **Description**: The dataset includes the HOUSE and TECH extensions of the SkillSpan Dataset. In the original work by Decorte et al., the test and development entities of the SkillSpan Dataset were annotated into the ESCO model.
* **License**: MIT, Please refer to the original source.
* **Modifications**: The datasets were used as provided without further modifications.

#### Qualification Mapping

* **Location**: inference/files/eval/qualification\_mapping.csv
* **Source**: Extended from the [Green Benchmark](https://github.com/acp19tag/skill-extraction-dataset) Qualifications
* **Description**: This dataset maps the Green Benchmark Qualifications to the appropriate EQF levels. Two annotators tagged the qualifications, resulting in a Cohen's Kappa agreement of 0.45, indicating moderate agreement.
* **License**: Creative Commons Attribution 4.0 International see Data License for details.
* **Modifications**: Extended the dataset to include EQF level mappings and the annotations were verified by two annotators.

#### Access and Usage

To use these datasets, ensure you comply with the original dataset's license and terms of use. Any modifications made should be documented and attributed appropriately to your project.

{% hint style="info" %}
For datasets requiring access tokens, such as those from HuggingFace 🤗, please contact the maintainers.
{% endhint %}


# Training

Train your entity extraction model using PyTorch.

First, activate the virtual environment as explained [here](/our-tech-stack/livelihoods-classifier/getting-started#dep).

### Train an Entity Extraction Model

Configure the necessary hyperparameters in the config.json file. The defaults are:

```json
{
    "model_name": "bert-base-cased",
    "crf": false,
    "dataset_path": "tabiya/job_ner_dataset",   
    "label_list": ["O", "B-Skill", "B-Qualification", "I-Domain", "I-Experience", "I-Qualification", "B-Occupation", "B-Domain", "I-Occupation", "I-Skill", "B-Experience"],
    "model_max_length": 128,
    "batch_size": 32,
    "learning_rate": 1e-4,
    "epochs": 4,
    "weight_decay": 0.01,
    "save": false,
    "output_path": "bert_job_ner"
}
```

To train the model, run the following script in the `train` directory:

```sh
python train.py
```

The training script is based on the [official HuggingFace token classification tutorial](https://huggingface.co/docs/transformers/en/tasks/token_classification).

### Train an Entity Similarity Model

Configure the necessary hyperparameters in the `sbert_train` function in the sbert\_train.py file:

```python
sbert_train(model_id='all-MiniLM-L6-v2', dataset_path='your/dataset/path', output_path='your/output/path')
```

To train the similarity model, run the following script in the `train` directory:

```sh
python sbert_train.py
```

The dataset should be formatted as a CSV file with two columns, such as 'title' and 'esco\_label', where each row contains a pair of related textual data points to be used during the training process. Make sure there are no missing values in your dataset to ensure successful training of the model. Here's an example of how your CSV file might look:

| title                   | esco\_label                 |
| ----------------------- | --------------------------- |
| Senior Conflict Manager | public institution director |
| etc                     | etc                         |

More information can be found [here](/our-tech-stack/livelihoods-classifier/datasets#entity-similarity).


# Advanced Topics

In this page we aim to give further details about the classes and functions located to the GItHub repository.

## inference/linker.`py`

### class EntityLinker

Creates a pipeline of an entity recognition transformer and a sentence transformer for embedding text.

#### Initialization Parameters

entity\_model : str, default='tabiya/roberta-base-job-ner' Path to a pre-trained `AutoModelForTokenClassification` model or an `AutoModelCrfForNer` model. This model is used for entity recognition within the input text.

similarity\_model : str, default='all-MiniLM-L6-v2' Path or name of a sentence transformer model used for embedding text. The sentence transformer is used to compute embeddings for the extracted entities and the reference sets. The model 'all-mpnet-base-v2' is available but not in cache, so it should be used with the parameter `from_cache=False` at least the first time.

crf : bool, default=False A flag to indicate whether to use an `AutoModelCrfForNer` model instead of a standard `AutoModelForTokenClassification`. `CRF` (Conditional Random Field) models are used when the task requires sequential predictions with dependencies between the outputs.

evaluation\_mode : bool, default=False If set to `True`, the linker will return the cosine similarity scores between the embeddings. This mode is useful for evaluating the quality of the linkages.

k : int, default=32 Specifies the number of items to retrieve from the reference sets. This parameter limits the number of top matches to consider when linking entities.

from\_cache : bool, default=True If set to `True`, the precomputed embeddings are loaded from cache to save time. If set to `False`, the embeddings are computed on-the-fly, which requires GPU access for efficiency and can be time-consuming.

output\_format : str, default='occupation' Specifies the format of the output for occupations, either `occupation`, `preffered_label`, `esco_code`, `uuid` or `all` to get all the columns. The `uuid` is also available for the skills.

#### Calling Parameters

text : str An arbitrary job vacancy-related string.

linking : bool, default=True Specify whether the model performs the entity linking to the taxonomy.

### class FrenchEntityLinker

French version of the entity linker. In order to use, we need to rewrite the reference databases to the French version of ESCO.

## `inference/evaluator.py`

## class Evaluator(EntityLinker)

Evaluator class that inherits the Entity Linker. It computes the queries, corpus, inverted corpus and relevant docs for the [InformationRetrievalEvaluator](https://github.com/UKPLab/sentence-transformers/blob/master/sentence_transformers/evaluation/InformationRetrievalEvaluator.py), performs entity linking and computes the Information Retrieval Metrics.

### Initialization Parameters

entity\_type: str Occupation, Skill, or Qualification to determine the exact evaluation set to be used.

## `util/transformersCRF.py`

### class CRF(nn.Module)

Implemented from [here](https://github.com/lonePatient/BERT-NER-Pytorch/tree/master).

A class that creates a linear Conditional Random Field model.

### class AutoModelForCrfPretrainedConfig(PretrainedConfig)

Configuration class that inherits from [PretrainedConfig ](https://huggingface.co/docs/transformers/en/main_classes/configuration#transformers.PretrainedConfig)HuggingFace class.

### class AutoModelCrfForNer(PreTrainedModel)

A general class that inherits from [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) class. The model\_type is detected automatically.

model\_type: str Possible options include `BertCrfForNer`, `RobertaCrfForNer` and `DebertaCrfForNer.`

### class BERT\_CRF\_Config(PretrainedConfig)

Custom class used for configuring BERT for CRF.

### class BertCrfForNer(PreTrainedModel)

BERT-based CRF model that inherits from [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) class.

Same as [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel).

#### Forward Parameters

Same as [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) except for

`special_tokens_mask` default: None. We use this option from HuggingFace as a small hack to implement the special\_mask needed for CRF.

### class ROBERTA\_CRF\_Config(PretrainedConfig)

Custom class used for configuring RoBERTa for CRF.

### class RobertaCrfForNer(PreTrainedModel)

RoBERTa-based CRF model that inherits from [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) class.

Same as [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel).

#### Forward Parameters

Same as [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) except for

`special_tokens_mask` default: None. We use this option from HuggingFace as a small hack to implement the special\_mask needed for CRF.

### class DEBERTA\_CRF\_Config(PretrainedConfig)

Custom class used for configuring RoBERTa for CRF.

### class DebertaCrfForNer(PreTrainedModel)

RoBERTa-based CRF model that inherits from [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) class.

Same as [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel).

#### Forward Parameters

Same as [PreTrainedModel HuggingFace](https://huggingface.co/docs/transformers/en/main_classes/model#transformers.PreTrainedModel) except for

`special_tokens_mask` default: None. We use this option from HuggingFace as a small hack to implement the special\_mask needed for CRF.

## `util/utilfunctions.py`

### class Config

Configuration class for the [training hyperparameters](/our-tech-stack/livelihoods-classifier/training#train-an-entity-extraction-model).

### class CPU\_Unpickler

A class that loads the tensors in the CPU.


# Contributing Guide

If you encounter any bugs or want to contribute to this project please use the following templates to [open an issue on GitHub](https://github.com/tabiya-tech/tabiya-livelihoods-classifier/issues).

## Bug report template

<table><thead><tr><th>name</th><th width="201">about</th><th width="92">title</th><th width="79">labels</th><th>assignees</th></tr></thead><tbody><tr><td>Bug report</td><td>Create a report to help us improve</td><td></td><td></td><td>ApostolosBenisis</td></tr></tbody></table>

**Describe the bug** A clear and concise description of what the bug is.

**To Reproduce** Steps to reproduce the behavior:

1. Go to '...'
2. Click on '....'
3. Scroll down to '....'
4. See error

**Expected behavior** A clear and concise description of what you expected to happen.

**Screenshots** If applicable, add screenshots to help explain your problem.

**Desktop (please complete the following information):**

* OS: \[e.g. iOS]
* Browser \[e.g. chrome, safari]
* Version \[e.g. 22]

**Smartphone (please complete the following information):**

* Device: \[e.g. iPhone6]
* OS: \[e.g. iOS8.1]
* Browser \[e.g. stock browser, safari]
* Version \[e.g. 22]

**Additional context** Add any other context about the problem here.

## Add feature template

<table><thead><tr><th width="187">name</th><th width="273">about</th><th width="95">title</th><th width="88">labels</th><th>assignees</th></tr></thead><tbody><tr><td>Feature request</td><td>Suggest an idea for this project</td><td></td><td></td><td></td></tr></tbody></table>

**Is your feature request related to a problem? Please describe.** A clear and concise description of what the problem is. Ex. I'm always frustrated when \[...]

**Describe the solution you'd like** A clear and concise description of what you want to happen.

**Describe alternatives you've considered** A clear and concise description of any alternative solutions or features you've considered.

**Additional context** Add any other context or screenshots about the feature request here.


# FAQs

#### **General Usage**

**1. What is the Tabiya Livelihoods Classifier?**\
The Tabiya Livelihoods Classifier is a tool that leverages advanced transformer-based neural networks to extract and categorize key entities from job descriptions. It supports tasks like occupation and skill classification using frameworks like ESCO.

**2. Who can benefit from using this tool?**\
It is designed for HR professionals, recruiters, career advisors, labor market researchers, and developers working on job-matching technologies or workforce analytics.

**3. What types of entities can the tool extract?**\
The classifier identifies and categorizes five entity types: Occupation, Skill, Qualification, Experience, and Domain.

**4. Is this tool compatible with any specific standards or frameworks?**\
Yes, it retrieves ESCO-related entries for Occupations and Skills, aligning with widely used European job classification systems. With minimal work, other taxonomies like O\*Net, could be integrated.

***

#### **Technical Functionality**

**5. How does the Tabiya Livelihoods Classifier work?**\
The process involves two main steps:

1. **Entity Extraction**: Identifies relevant entities in job descriptions.
2. **Similarity Vector Search**: Matches extracted entities to entries in pre-defined frameworks or datasets.

**6. Does the tool use machine learning models?**\
Yes, it utilizes transformer-based models, which represent the state-of-the-art in natural language processing.

**7. Can I customize the classifier for specific industries or datasets?**\
The tool supports customization, allowing users to adapt the similarity search or integrate custom datasets to suit specific domains and use cases.

**8. What is the difference between entity extraction and similarity vector search?**\
Entity extraction identifies relevant entities from text, such as a job title or skill. Similarity vector search then matches these entities to related entries in a knowledge base, like ESCO, for standardization.

**9. Where can I find more technical details about the methods used?**\
For a deeper explanation of the underlying techniques, model design, and evaluation methodology, see the [research paper](https://arxiv.org/abs/2512.03195) *Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies* (Saroglou *et al*., 2025).

***

#### **Integration and Setup**

**10. How do I install and use the classifier?**\
Detailed installation and setup instructions are available in the [user guide](/our-tech-stack/livelihoods-classifier/getting-started#quickstart-guide).

**11. Can the classifier be integrated into existing HR systems?**\
Yes, it is designed to be easily integrated into workflows or systems through APIs or library functions.

**12. Are there any prerequisites for using this tool?**\
A working knowledge of Python is recommended for setup and integration. Familiarity with natural language processing concepts is beneficial but not mandatory.

***

#### **Performance and Limitations**

**13. How accurate is the entity classification?**\
The classifier achieves state-of-the-art entity recognition results based on the dataset released by [Green et al](https://aclanthology.org/2022.lrec-1.128/). Albeit, as with any machine learning model, the Entity Linker is not perfect. If you encounter bugs or inappropriate use-cases, please open an issue on [GitHub](https://github.com/tabiya-tech/tabiya-livelihoods-classifier/issues)!

**14. Does the tool handle multilingual job descriptions?**\
We are currently working on developing a method of expanding the tool's capabilities for multiple languages. As of right now, the tool supports only the English and French languages.

**15. Are there limitations on the size of input text?**\
The model uses the [NLTK sentence tokenizer](https://www.nltk.org/api/nltk.tokenize.sent_tokenize.html) function to handle large texts, so theoretically, there is no limit to the input text size. In the current version, the BERT-based models used for entity extraction have a limit of 128 tokens (roughly 100 words). You can use the[ training script](/our-tech-stack/livelihoods-classifier/training#train-an-entity-extraction-model) to retrain the model to fit your needs.

***

#### **Support and Customization**

**16. Can I contribute to or extend the tool?**\
Yes, developers are welcome to customize and extend the tool. Refer to the contributing guide in the documentation for guidelines.

**17. Where can I find support or report issues?**\
Support is available through the official repository or customer service channels. Issues can be reported on the [GitHub issues page](https://github.com/tabiya-tech/tabiya-livelihoods-classifier/issues) or via email.

**18. Are updates and new features planned?**\
Yes, the tool is actively maintained, with plans for additional features and improved integrations based on user feedback.


# Compass

Unlock your potential, discover your skills

<figure><picture><source srcset="/files/gBQ8pMjIiS3x0YOqsrxD" media="(prefers-color-scheme: dark)"><img src="/files/CuJIh9eZLRkgJWPkeCK7" alt="Compass by Tabiya logo" width="339"></picture><figcaption></figcaption></figure>

Compass is an innovative, AI-powered chatbot designed to revolutionize the way a young person identifies, articulates, and showcases their skills. Developed by Tabiya, Compass is a personal career assistant, helping youth uncover hidden talents and match them with the best opportunities in the job market.

## The Compass Solution

Compass is an open-source, conversational AI tool that:

1. Engages in **natural dialogue to explore** your experience
2. Analyzes a user's input to **identify and categorize skills** against localized [taxonomies](/our-tech-stack/inclusive-livelihoods-taxonomy/open-taxonomy-platform)
3. Identifies **skills from both formal and informal work**
4. Creates **a comprehensive skills profile** tailored to the user
5. Generates a **professional customizable CV** highlighting one’s strengths
6. **Matches a user's skills** with relevant economic opportunities
7. Recommends **personalized recommendations** for skill development and career advancement.

Compass uses a large language model (LLM) and a conversational interface to help job seekers build a CV that highlights their skills. The tool combines a commercial LLM with a human-reviewed skills taxonomy for the labor market.

## Addressing Critical Challenges

In today's rapidly evolving job market, Compass tackles two persistent and interconnected problems that hinder effective workforce development:

* **Challenges in showcasing skills:** Crafting a CV that highlights relevant skills is tough for many job seekers. It's about translating experiences into terms that appeal to employers. This is harder for those with non-traditional or informal experience. A poor CV can lead to missed opportunities.
* **The scalability struggle:** Traditional methods rely heavily on human career counselors. While these professionals offer valuable insights, this approach has significant limitations.

Compass breaks down barriers by offering an AI solution that's scalable, affordable, and high-quality.

## Benefits of Compass

<table data-view="cards"><thead><tr><th></th><th></th><th></th></tr></thead><tbody><tr><td><em><strong>For Partners</strong></em></td><td><ul><li><strong>Increase Efficiency:</strong> Simplify skills identification and job matching.</li><li><strong>Improve Outcomes:</strong> Enhance job placements and retention with suitable opportunities.</li><li><strong>Cost-Effective Scaling:</strong> Offer personalized guidance to more job seekers without extra staff and allow consellors to focus on advising jobseekers.</li><li><strong>Data-Informed Decisions:</strong> Use insights to tailor services and programs.</li></ul></td><td></td></tr><tr><td><em><strong>For Job Seekers</strong></em></td><td><ul><li><strong>Discover Potential:</strong> Identify and articulate hidden skills.</li><li><strong>Access Guidance:</strong> Benefit from AI-driven career advice with human expertise.</li><li><strong>Improved Matching:</strong> Find opportunities that fit unique skill sets.</li><li><strong>Career Development:</strong> Receive tailored recommendations for skill and career growth.</li></ul></td><td></td></tr><tr><td><p><em><strong>For Funders</strong></em></p><ul><li><strong>Scalable Impact:</strong> Reach thousands of job seekers efficiently with minimal cost increase.</li><li><strong>Data-Driven Insights:</strong> Provide valuable data on skills gaps and labor market trends to guide policies and programs.</li><li><strong>Promote Equity:</strong> Value skills from diverse backgrounds, including informal and unpaid work.</li><li><strong>Enhance Existing Programs:</strong> Strengthen current workforce development initiatives.</li><li><strong>Foster Innovation:</strong> Use the Tabiya ecosystem to encourage innovation and address social challenges.</li></ul></td><td></td><td></td></tr></tbody></table>

## Upcoming Features

Our vision for future features and our roadmap [can be found here](/our-tech-stack/compass/roadmap).

## Our Open Source Commitment

Compass is designed as a digital public good:

* The core technology is open-source, allowing for transparency and community-driven improvements
* We aim to build a diverse global community of contributors and implementers around Compass
* Organizations supporting youth in their career journeys can adapt Compass to their specific needs and contexts

## Get Involved

We welcome partnerships and collaborations to further develop and implement Compass:

* For inquiries about supporting or implementing Compass, connect with us [here](https://go.tabiya.org/contact)
* Follow our progress and [join the conversation on Linkedin](https://www.linkedin.com/company/tabiya)

Together, we can leverage the power of AI to create more inclusive and efficient labor markets worldwide.

## Funders and Partners

<figure><img src="/files/zT0hO5M8kTRPV9c7xcOd" alt=""><figcaption></figcaption></figure>


# Technical Overview

### AI Architecture

Compass utilizes agentic workflows to interact with users, gather information, and identify their skills.

Compass mimics how a human would approach a conversation and the resulting tasks, acting as an overarching agent that decomposes into smaller agents, each with its own responsibilities and goals.

Each agent within Compass has a specific responsibility and performs multiple tasks to achieve its goal. For example, an agent might converse with a user to collect specific information, process that information, and prepare it for use by another agent.

<figure><img src="/files/6fAVkRmiecWFeSCxiL1l" alt=""><figcaption><p>Compas AI architecture overview</p></figcaption></figure>

Agents maintain an internal state that allows them to apply a strategy to accomplish their goal, and have access to the user's conversation history and use tools based on LLM prompts (or not). These tools could be used for tasks such as conversing with the user, named entity extraction, classification, or transforming user input.

Agents are guided by a combination of instructions (prompts) and their internal state. The prompts can vary based on the agent's state to help them accomplish their tasks.

Once Compass has gathered all the necessary information from the user, it processes the data to identify the user’s top skills. This identification is done through a multi-stage pipeline that employs various techniques, such as clustering, classification, and entity linking the a occupations/skills taxonomy.

<figure><img src="/files/Cahjp2iFVpUk9X9Im663" alt=""><figcaption><p>Detailed AI architecture with the multi-stage skills pipeline</p></figcaption></figure>

Compass leverages LLMs in four ways:

1. **Conversational Engagement**: Unlike typical applications where the LLM responds to user questions, Compass reverses this interaction. It generates questions to guide a **directed** and **grounded** conversation.
2. **Natural Language Processing Tasks**: Compass uses the LLM for tasks like clustering, named entity extraction, and classification, handling user inputs efficiently without the need for costly and time-consuming model training or fine-tuning.
3. **Explainability and Traceability**: The LLM provides reasoning for specific outputs, allowing for explanations that link discovered skills back to the user’s input. This feature, which is based on a variation of Chain of Thought reasoning, is especially noteworthy—not only for the capability it offers but also because it was an unplanned outcome. It emerged while attempting to solve a different problem: improving the accuracy of the LLM's tasks
4. **Filtering of Taxonomy output**: The LLM filters relevant skills and occupations connected to the ESCO model from the conversation's output. Leveraging its advanced reasoning capabilities, the LLM efficiently processes large amounts of text, identifying the most pertinent entities. This approach combines traditional entity linking via semantic search with LLM-based filtering, creating a hybrid solution.

Compass is grounded and protected from hallucinations in multiple ways:

* **Task Decomposition**: Smaller, more manageable tasks are assigned to individual agents with specific LLM prompts.
* **State-Induced Instructions**: Agents use their internal state to generate targeted instructions during user interactions, guiding the conversation toward a specific goal. This approach reduces the size of the prompt by including only relevant segments, making it more likely that the LLM will follow the instructions accurately, thereby reducing the risk of hallucinations.
* **Guided Output**: Instructions are carefully crafted to increase the likelihood of relevant responses. Techniques include:
  * One- and few-shot learning
  * Chain of Thought
  * Retrieval Augmented Generation
  * JSON schemas with validation and retries
  * Ordering output segments to align with semantic dependencies
* **State Guardrails**: Simple, rule-based decisions are made whenever possible, reducing reliance on the LLM and minimizing potential inaccuracies.
* **Taxonomy Grounding**: By linking entities to a predefined occupations/skills taxonomy, Compass ensures that identified skills remain within a relevant and accurate domain.

### Evaluation

For evaluating Compass, we followed the strategy outlined below:

* **Rigorous Embeddings Evaluation**: We rigorously evaluated various strategies for generating embeddings from the taxonomy entities. Our considerations included identifying which properties of the entities should be included in the embeddings generation, as well as determining the optimal number of entities to balance accuracy and precision. For the tests, we used established datasets from the literature and generated synthetic data to mimic Compass user queries.
* **Isolated Component Testing**: Each agent's tools were evaluated individually using specific inputs and expected outputs. For example, classification components were tested with known inputs to verify accurate label assignments.
* **Scripted Conversations**: Conversational agents were tested using predefined dialogues, with outputs evaluated by either automated evaluators (other LLMs) or human inspectors.
* **Simulated User Interactions**: Compass was tested in end-to-end scenarios by simulating user interactions driven by an LLM. The simulated user was given a persona based on our UX research and additional instructions to cover specific cases of interest. These conversations were then assessed for quality and relevance by automated evaluators (other LLMs) or human inspectors.
* **User Testing and Trace Analysis**: On a smaller scale, real user tests were conducted. By tracing the top skill outputs back to the user's input, human inspectors could assess the performance of specific agents within Compass.

### The Core Role of a Taxonomy

Tabiya's inclusive taxonomy plays a central role in Compass. It grounds the LLM’s tasks, but there are several additional aspects worth mentioning:

* **Standardization**: The identified skills are linked to a standard taxonomy, making interpretation and comparison easier. The concepts behind these skills are well-defined and can be explained, allowing for clarification and disambiguation.
* **Canonicalization**: Explored skills are listed with canonical names and UUIDs, enabling consistent referencing across different experiences and applications.
* **Network Structure**: The taxonomy models the labor market by associating occupations with skills, forming a knowledge graph that can provide additional insights to users.
* **Unseen Economy**: The taxonomy has been extended to include activities from the unseen economy, empowering young women and first-time job seekers to enter the job market.
* **Localization**: A taxonomy can consider the specific context of a country. This includes occupations unique to certain regions, alternative names for occupations that are region-specific, and varying skill requirements for the same occupation across different countries.
* **Work Type Classification**: All experiences are classified into four types (wage employment, self-employment, unpaid trainee work, and unseen work), which allows for a more targeted exploration of the job seeker’s skills.

### Technical Stack Overview <a href="#technical-stack-overview" id="technical-stack-overview"></a>

* **Language Models and Embeddings**: Compass utilizes the `gemini-2.0-flash-001` model for its LLM capabilities and the `text-embedding-005` model for embeddings. The `gemini-2.5-pro-preview-05-06` model is used for the LLM auto-evaluator. The Gemini model was chosen for its balanced performance across task accuracy, inference speed, rate availability, and cost.
* **Backend Technologies**: Developed with `Python 3.11`, `FastAPI 0.111`, and `Pydantic 2.7` for a performant server-side environment. An asynchronous framework suited the use case well, as LLM inference endpoints can be slow. Python was chosen for its extensive AI/ML library support and because it made it easier to integrate ML scientists into the development team.
* **Frontend Technologies**: The UI, built with `React.js 19`, `TypeScript 5`, and `Material UI 5`, is optimized for mobile but performs well on tablets and desktops. Additionally, we use `Storybook 8.1` to showcase, visually inspect, and test UI components in isolation.
* **Data Persistence**: Data is securely stored using `MongoDB Atlas`, which includes vector search capabilities. Our team was already familiar with `MongoDB`, and the taxonomy was already in `MongoDB Atlas`, so it was a natural choice.
* **Deployment**: The entire application is deployed on `Google Cloud Platform (GCP)`, ensuring high availability and scalability. We use `Pulumi` to deploy nearly all the infrastructure, as it allows us to write deployment code in `Python`, aligning with the rest of the backend development. Additionally, our team was already experienced with `Pulumi`, making it a natural choice. For error tracking and application performance monitoring, we use `Sentry`.

<figure><img src="/files/KeIn6ByyUp4kohp3NLGU" alt=""><figcaption><p>Cloud architecture</p></figcaption></figure>


# Compass API

Compass is an innovative, AI-powered chatbot designed to revolutionize the way a young person identifies, articulates, and showcases their skills. Learn more about the API that powers Compass.

## Overview and Format

The [Compass](/our-tech-stack/compass) API provides secure access to users' preferences, conversation state, and experiences and skills, enabling seamless integration with your applications. The base URL for all API requests is <https://demo.compass.tabiya.tech/api/docs>. To ensure standardized communication, all requests and responses (except file uploads) utilize the JSON format, making it straightforward to integrate with any modern programming environment.

## Credentials and Authentication

Before integration, developers must obtain the necessary credentials.&#x20;

* API Keys: An API Key, can be requested by contacting the admins via the designated email. Once issued, the API key must be included in the x-api-key header of every request. For example, the header should be formatted as: x-api-key: \<your-access-token>.
* Authorization Headers: These are compass user-specific and require first logging in via the compass app to obtain the access token.

## Direct Access to the OpenAPI Specification

Two access methods are available for the API specifications. The primary method is through the interactive Swagger UI documentation page, which allows browsing and live testing of all endpoints. Crucially, for developers looking to quickly configure their clients, a direct, standalone link to the OpenAPI v3 JSON specification file is provided. This link is essential for developers who want to import the entire API definition into tools like Postman or Insomnia (using the *Import > Link option*), accelerating setup by automatically generating all endpoint, parameter, and schema definitions:

* OpenAPI Specification link: <https://demo.compass.tabiya.tech/openapi.json>
* Swagger link: <https://demo.compass.tabiya.tech/api/docs>

## Usage Policies and Rate Limiting

To ensure fair and stable access for all users, we enforce a Rate Limiting policy. Developers are limited to 120 requests per minute per unique API key. Exceeding this threshold will immediately result in an HTTP 429 "Too Many Requests" status code being returned, and subsequent requests will be blocked until the minute resets.


# UX Evaluation

In August 2024, Tabiya tested the user experience of Compass with job-seekers recruited from Harambee Youth Employment Accelerator.

## Background and Motivation

Tabiya, in partnership with Harambee Youth Employment Accelerator, conducted a three-day series of user experience (UX) tests devised to answer the following questions:

1. Can participants initiate and maintain a natural conversation with Compass?
2. Does Compass effectively guide users through the skill identification process?
3. Do users find the skills identified by Compass relevant and accurate?
4. Are participants satisfied with the overall experience and feel they have a clearer understanding of their skills?

[In the South African case, the envisaged users of Compass are the country’s youth, who face high and rising disengagement in the labor market, education, and training rates](#user-content-fn-1)[^1]. This is due, in part, to[ job search distortions](#user-content-fn-2)[^2]. That is, work-seekers may have incomplete information about their skills. This means that work-seekers could apply for jobs that they are incompatible with or [stop searching altogether](#user-content-fn-3)[^3]. Compass is built to mitigate this particular search distortion, which forms part of the drivers of unemployment, especially amongst the youth.

## Methodology

South African youth (between the ages of 18 and 35) subscribed to Harambee’s SA Youth Mobi Platform were shown an advertisement to join a UX testing session hosted at Harambee offices in Cape Town. Upon receipt, 604 were deemed most compatible with the selection criteria. This criterion includes the following:

1. Young person is subscribed to SA Youth Mobi Platform
2. Young person ordinarily resides no more than roughly 15 kilometeres away from the Harambee Cape Town office
3. Young person has had work experience before
4. Young person is actively looking for employment

From this criterion, black female participants and participants who had evidence of starting their own small business were favored in the final UX test participant selection.

Before the UX test session begins, each participant is given a consent form to read and sign. Where applicable, participants were asked to sign a photography consent form if they were comfortable doing so. Following this, each UX test comprised four parts. The first part consisted of an introduction and asking the participant for permission to record the session. This part of the UX testing session was also used to explain what Compass is and how it works to the participant.

The second part consisted of the commencement of the UX test. During this part, each participant interacted with Compass as independently as possible until their experiences and skills were summarized, or until the moderator indicated they should stop interacting with Compass due to a time constraint. The UX test involved Compass gathering information from participants on their work experiences, including paid and unpaid work in establishments or non-establishment settings. This is referred to as the work experience section throughout this report. Following this, Compass produced an experience summary listing the participants' experiences, locations, and durations chronologically. Next, Compass gathered information on what skills participants had obtained from each work experience through an interactive dialogue. Up to five skills per experience were subsequently listed at the end of the dialogue. This is referred to as the skills assignment section throughout this report.

During the third part of the session, the moderator asked the participant a series of post-test feedback questions. Finally, the moderator closed the session by giving each participant the timeframe to expect their monetary incentive (of R450) for participating in the UX test. The closing was also used to ask participants if they were comfortable with being placed in a Harambee WhatsApp group for an expert panel that is periodically invited to participate in Harambee sessions targeted at its youth database. All four parts were projected to take 45 minutes but could take up to one hour. In addition to R450, each participant was given a snack pack before each UX test. This is done to ensure that participants had food to eat before commencing their UX tests. The complete discussion guide for each UX test can be found [here](/our-tech-stack/compass/ux-evaluation/ux-testing-discussion-guide).

Compass was designed to extract information from its users conversationally. As such, Compass functions similarly to a messaging application. Compass textually asks its users for information pertaining to their experiences and allows users to respond before concluding or probing for more information. Given that this exercise operates similarly to an interaction where someone would communicate through text messaging on their smartphone, the UX test was conducted on a cellphone. For heightened accessibility and ease of use, the cellphone chosen was the Samsung Galaxy A23 because this is the most widely used smartphone among the youth in Harambee’s database.

## Results

1. Demographics

A set of 11 SA Youth subscribed youth were selected to test the functionality of Compass. Three of these participants tested out Compass on the 27th of August. Two were male, and one was female. Four were scheduled to complete the testing on the 28th, but there was an attrition rate of 50% on this day. A male and a female showed up; the two not in attendance were male. The 29th of August was the final day of testing. Four participants, three female and one male, each tested Compass. Therefore, 82% (9 people) of the chosen cohort of 11 people participated in the UX Testing. 44% of the UX testers participants were male and 66% female.

2. Timing

From the nine young people who tested Compass, two groups emerged. Group 1 comprised three participants who received a work experience summary but not a skills summary. Group 2 included six participants who received both work experience and skills summary. Group 1 took an average of 32 minutes to complete the entire UX testing session. Group 2 took an average of 32 minutes and 3o seconds to complete. Overviews of each UX testing session, including details on the timings of each session can be found [here](https://github.com/tabiya-tech/docs/blob/main/projects/compass/ux-evaluation/broken-reference/README.md).

The full Compass experience was completed without interruption only once. This was done by Participant 9 who only possessed one experience and thus, one set of skills. Hence, the initially projected time of 45 minutes to complete the entire UX testing session was an underestimation. Instead, all participants finished within 45 minutes only because the moderator of the participant sessions prematurely instructed participants to finish their interactions with Compass. Given that participants were paid for their time, not doing this would have violated labor laws.

3. Participant Feedback

Participants responded extremely favourably towards Compass. 88.9% of participants reported that they found Compass easy to use and understand its responses. This feedback suggests that Compass allowed for the initiating and maintaining of natural conversation across this group of UX testers. In fact, participants were particularly impressed with the fact that they could have conversations with Compass “as if it were a human”. This aided in their ability to independently complete their work experience and skills obtaining journeys. 77.8% of participants found that Compass identified accurate and relevant skills without omitting any irrelevant skills. 88.9% of participants said that Compass helped them gain a better understanding of their experiences and skills, particularly because of how simply it put these in their summaries.

All participants welcomed the prospect of Compass suggesting jobs and sectors for them to apply to given their experiences. Additional feedback included having personal traits and characteristics as part of Compass skills summaries and including qualifications in the Compass CV. 100% of the participants who completed the UX test said they would recommend Compass to other job seekers. All of this taken together is strong evidence in favor of the hypothesis that participants are satisfied with their overall experience using Compass.

4. Positive Observations

**Predictability**. Over time, participants began to anticipate the flow of questions asked by Compass for each experience. This yielded quicker responses since participants could combine answers to two separate questions that they eventually knew would be asked by Compass (for example, what the experience was and where it was) without being prompted to do so.

**Persistence**. Compass’ persistence, particularly in the unseen economy, is a strength. Compass asks participants a series of headline questions to initially characterize their experience in the seen or unseen economy. After an experience is imputed by a participant, Compass repeats the question that has been answered before moving onto the next work experience category. This repetition of questions allows participants to note all of their work experience (and, in turn, their associated skills), within both the seen and unseen economy, which is particularly important for the unseen economy.

**Agility**. Throughout participants’ user journeys, Compass was able to effectively redirect the conversation back to topic when participants misunderstood questions. Similarly, Compass could navigate through most grammar and spelling mistakes, which each participant made at least one of. The only spelling mistake Compass did not automatically rectify was “Checkers”, which was erroneously spelt as “Checker” by Participant 4. This spelling error persisted until the experience summary was presented to the user. This notwithstanding, participants have several opportunities to correct spelling errors through their Compass user journeys.

Another instance of Compass’ agility is in the case of Participant 6, who incorrectly imputed a volunteer role as a paid role. When the participant responded to a question that asked if she had worked for a company or business for money, Compass identified key words, including “volunteer”, in her response and then proceeded to probe into whether her imputed experience was indeed paid or unpaid. After uncovering that it was an unpaid position, Compass was able to accurately classify this role as such and correctly list it Participant 6’s work experience summary.

5. Considerations

**Misunderstandings**. While Compass has yielded objectively great results across this cohort, it would be remiss to omit elements to consider when rolling it out to larger audiences. The first consideration is the language barrier effect. Compass requires a moderate proficiency in English. For the South African youth case, English is likely not the average prospective Compass users’ first language.

Each participant in the Tabiya-Harambee UX testing session possessed an average proficiency level high enough to get through the UX test and the questionnaire afterwards. However, there were times, such as in Participant 4’s UX test, when the moderator and participant conversed in the participant’s native language when the participant asked questions or when the moderator explained what Compass is. Moreover, it was clear that Compass posed some questions that caused some confusion, misunderstanding or hesitation among participants. Some examples include:

1. “Can you tell me about the first experience you had working for yourself?” To this question, Participant 1 described the independence and fulfilment he derived from this experience rather than describing the experience itself. Compass redirected the conversation by subsequently asking the participant what kind of work he participated in.
2. “...what was a typical day like at work” is the phrase used by Compass to initiate information on what skills participants obtained from their various work experiences. Participant 1 provided information related to the atmosphere or external happenings of a typical day rather than the tasks that they completed on a typical day. Compass redirected the conversation by asking what tasks participants were responsible for.
3. When Compass asked Participant 5 to “tell me what your experience was like”, the participant responded, “It was a good experience to volunteer”, rather than going into detail about what tasks the experience entailed.

Compass’ ability to infer findings and redirect conversations despite misinterpretations is a glaring strength in this regard. However, given that only 11% of participants completed Compass’ experiences and skills assignment sections in time, refining misunderstood questions to save time may be prudent. Compass’ ambition to expand its language offerings beyond English also bodes well for mitigating misunderstandings.

**Confidentiality**. A related but slightly different consideration worth mentioning is how Compass uses erroneously imputed, confidential information. Throughout each of the interviews conducted, Compass did not directly ask for sensitive or confidential information. However, there were times throughout the UX test where participants misinterpreted or erroneously answered one of Compass’ questions. In Participant 3’s case, his misunderstanding of the question “Can you tell me, was this a paid job?” caused him to impute the exact salary he earned from the job, rather than to affirm that the position was paid. This information did not appear on the Participant’s CV; therefore, it had no impact on his results for this version of Compass. However, if the ambition is to eventually roll out a version of Compass that uses the conversations it has with young people to suggest suitable jobs and sectors to find potential jobs, then this could be an issue. Compass could, in this case, provide a participant who has erroneously imputed their exact salary with job opportunities that pay within that same salary range and, in so doing, narrow their job application options relative to a participant whose remuneration is not known by Compass.

**Skills Reporting**. Participant 5 noted that “patience and accuracy” were skills required to do one of the roles he had well. Despite asking for this information, Compass omitted these skills from his skills summary. It is worth noting that Participant 5 felt that these skills were not as important as the skills surfaced by Compass. Conversely, Participant 7 noted that a great personality and good communication were essential for working at the company that he previously worked for. While neither of these were explicitly included in his skills summary, “customer service” - a plausible alternative to name these characteristics - was.

Whether or not to include skills explicitly mentioned by the participant is not easily answerable. For some participants, such as Participant 5, the skills summarized by Compass (which are based on the European Skills, Competences, Qualifications and Occupations (ESCO) framework) seem to resonate well with their experiences. In contrast, in other cases, such as the case of Participant 4, self-assigned personal characteristics and attributes would be a value addition in a CV. Nonetheless, if Compass explicitly asks participants to impute skills which they feel are important in their role, it seems worthwhile that Compass ought to, at the very least, find the most closely related ESCO skill to add to the participant’s skill summary.

In addition, when asked to provide information about her experience as a cashier in the work experience attribution section of Compass, Participant 4 noted that as a cashier, she was “assisting customers with electronic and cash handling payments, scanning and packing customers groceries, counting float and cash up”. Despite these detailed tasks and skills put forward by this participant, Compass did not include any of these in her work experiences summary. This is likely because she had not yet reached the portion of Compass that asks directly about tasks participants about their gained skills from a particular job. Unfortunately, Participant 4 did not proceed to the skills attribution section of Compass because of insufficient time to do so. This left Participant 4 feeling frustrated and disappointed. Ideally, Compass should be able to identify skills information given in the work experience section and preemptively add this skills information to a young person’s work experiences summary, so they do not have to repeat this information in the skills assignment section.

**Experience Reporting**. There are several instances where Compass erroneously summarizes participants' work experiences. For example, Participant 2’s unseen economy experience was duplicated in her work experience summary, and similarly, Participant 4’s cashier experience was erroneously duplicated.

One of Compass’ opening questions to Participant 6 was, “Have you ever worked for a company or someone else’s business for money?” Participant 6 responded to this question by saying, "Yes I am currently working for a company.” Compass then proceeded to gather information on the name and location of the company as well as how long she had worked there. However, Compass did not ask her what the title of her role was. Therefore, in her experiences summary, this work experience was named “Working for a Company”.

In another instance, Participant 5 erroneously classified waged employment as contract work (this is unsurprising since the term contract work could easily be interpreted as work done after signing a contract, especially if English is not a participant’s native language). During the dialogue, Compass rectified this error and correctly identified the participant’s experience as waged employment rather than contract work. However, Compass included “Contract work (Self-employment)” in the participant’s work experience summary despite this. Finally, Participant 5 worked remotely as a data capturer for a United Kingdom-based organization. When Compass summarized this experience, the fact that this experience was done remotely was omitted. Therefore, it may appear that the participant worked in the United Kingdom to a prospective employer.

**Overlapping Experiences**. Some South African youth hold more than one job at the same time to maximize their income. Often, one job may be held as a paid opportunity undertaken by an employer, while the other is a microentrepreneurship or freelance role that the young person independently does. Participant 1, for example, held a volunteer and a self-employment position while undertaking waged employment. Participant 2 mentioned that she would only ordinarily list some of the experience items she disclosed to Compass on her CV. It is useful to understand whether South African employers are more or less likely to hire someone with a CV that includes overlapping experiences relative to someone who performs one experience at a time. If an employer thought that overlapping self-employment (or volunteer) roles may create a conflict of interest or time mismanagement issue, even if this may not be the case, this could work against a jobseeker’s employment prospects. Alternatively, if an employer viewed holding multiple positions simultaneously as a signal of a hard-working nature, this could work in favor of a jobseeker. Whether a young person is allowed or encouraged to list overlapping full-time experiences ultimately ought to depend on the young person's objective and employees' expected responses.

**Unseen Economy Bottlenecks**. Compass is more susceptible to producing errors or yielding misunderstandings when it asks for unseen economy experiences relative to seen economy experiences. For example, Participant 2 was asked to give specific information about her experience with helping others; she said she was “well known in the community”. Compass, unsatisfied with this answer, then asked the question again in the same way. This led to the participant flagging this with the moderator. Ultimately, Compass erroneously duplicated the unseen economy experience provided by Participant 2 on her CV.

In another instance, when Participant 4 was asked whether she has “ever helped out friends or family members without getting paid”, she responded by describing an instance where she charitably gave her colleague money for transport and food. Compass failed to identify that this activity did not meet the requirements of an unseen economy activity and proceeded to gather more information to (erroneously) list this activity in the participant’s experience summary.

Participant 5’s unseen economy activity was taking his parents shopping. The Compass experience reporting structure urged the participant to provide a date when this activity happened. For the unseen economy, it can be particularly difficult to confine ongoing sporadic activities to a start date and end date. Indeed, it is plausible that Participant 5 took his parents shopping on several occasions despite reporting only the most recent date that he had done this activity. Moreover, Compass’ classification of the participant’s self-reported activity of taking is parents shopping as “helping out friends” in his work experience summary seems ill-defined.

The unseen economy is conceptually challenging to articulate non-technically. Moreover, assigning a “location, duration, organization name” structure to unseen economy may not work as seamlessly as in the seen economy. Hence, the current unseen economy prompts can do with revision in light of the above-mentioned bottlenecks.

**External Validity**. Finally, it is vital to understand the validity constraints of this exercise. These results hold for the participants who completed the UX test during the time they did so. However, due to the relatively small sample size of participants who completed the UX test, these results cannot be interpreted as applying to any other group.

## Conclusion

Tabiya worked with Harambee Youth Employment Accelerator, to roll out a Compass UX Testing Session to nine South African youth who made up the session’s participants. The participants who completed the UX test had diverse demographics. Each session was managed and moderated by a Harambee staff member. In addition to using Compass on a smartphone, participants answered a series of questions to gauge how they felt about their interaction with Compass.

The feedback provided by participants is in overwhelming favor of the ease of use, accuracy and relevance of Compass in identifying their skills and experiences. The participants who used Compass were largely able to do so independently. This positive feedback is encouraging and extremely compelling, but it must be balanced with the considerations that must be made before rolling out Compass.

[^1]: See Statistics South Africa. 2024. P0211 - Quarterly Labour Force Survey (QLFS). Statistical Release, Pretoria: Statistics South Africa.

[^2]: See Carranza, Eliana, Robert Garlick, Kate Orkin, and Neil Rankin. 2022. "Job Search and Hiring with Limited Information about Workseekers' Skills." American Economic Review, 112 (11): 3547–83.

[^3]: See Michèle Belot, Philipp Kircher, Paul Muller, Providing Advice to Jobseekers at Low Cost: An Experimental Study on Online Advice, The Review of Economic Studies, Volume 86, Issue 4, July 2019, Pages 1411–1447, [https://doi.org/10.1093/restud/rdy059.](https://doi.org/10.1093/restud/rdy059.See)

    See Conlon, John J., Laura Pilossoph, Matthew Wiswall, and Basit Zafar. Labor market search with imperfect information and learning. No. w24988. National Bureau of Economic Research, 2018.


# UX Testing Discussion Guide

## Introduction and Warm-up

Hi \[participant’s name], thank you so much for taking the time to join me today!

My name is \[moderator name], and I'm helping with a research project to understand how a new tool called Compass might be able to assist job seekers in their job search journey.

Compass uses research and AI to help job seekers explore their skills and provides guidance on how best to leverage them in their job search.

Your feedback will help us make improvements to Compass so it can work better for job seekers like yourself.

Participant's consent: did you get a chance to sign the consent form?

* Everything you share in this session will be kept confidential. We won't use your name or any identifying details in our reports.
* So please feel comfortable sharing openly and honestly
* Also, there are no right or wrong answers, the most helpful to us is to hear your perspective
* Please also keep everything you see and everything we talk about today confidential

Permission to record: I have a small request; would it be okay with you if I record our session?

* I’ll be engaged in the conversation with you and want to make sure I don’t miss anything important that you tell me, only myself and the team will be able to see it and we delete everything after the completion of the study

Permission to livestream: I have one more request; would it be okay with you if I livestream our session?

* Members of the team working on Compass are eager to learn from your experience and they would love to observe the session.

Thank you so much, much appreciated \[if they say yes]

No worries at all \[if they say no]

Do you have any questions for me?

Before we jump into trying out Compass, I'd love to hear a bit about your experience looking for a job.

* Can you briefly tell me a little about your job search so far?

Thank you for sharing a bit about your job search journey, let’s get started with Compass.

## Tasks

### Task 1: Initial Interaction \[10 min]

* Scenario \[Moderator to read this part]:
* You are a job seeker exploring finding a job. You want to find work that suits you, based on your skillset. Start a conversation with Compass to identify your skills.
* Instructions: \[Moderator to read this part] (to avoid bias, alternate between the two options as you go through the interviews)
* Option 1: how would you go about using Compass to identify your skills.
* Option 2: introduce yourself and tell Compass you are looking to identify your skills for your job search.
* Feedback: \[Moderator to read this part]:
* What do you think of the conversation flow so far?
* On a scale of 1 to 5, how satisfied are you with this conversation with Compass so far? (1 unhelpful 5 very helpful)
* Success criteria: participant successfully initiates a conversation with Compass.

### Task 2: Skill Identification \[10 min]

* Scenario \[Moderator to read this part]:
* As you can see, Compass has asked you questions about your various experiences
* Instructions: \[Moderator to read this part]:
* Go ahead and answer Compass's questions about your past roles and responsibilities to the best of your ability
* Feedback: \[Moderator to read this part]:
* What do you think of how Compass is trying to identify your skills?
* On a scale of 1 to 5, how satisfied are you with this conversation with Compass so far? (1 unhelpful 5 very helpful)
* Success criteria: participant answers questions about their work experience and engages in a dialogue with Compass

### Task 3: Skill Review & Feedback \[10 min]

* Scenario \[Moderator to read this part]:
* As you can see, Compass has generated a list of your potential skills
* Instructions: \[Moderator to read this part]:
* Go ahead and take a moment to review these skills identified by Compass
* Feedback: \[Moderator to read this part]:
* How accurate/relevant do you feel the skills Compass identified for you are? Why or why not?
* On a scale of 1 to 5, how helpful was Compass in identifying your skills? (1 unhelpful, 5 very helpful)
* On a scale of 1 to 5, how satisfied would you say you are with Compass’ ability to identify your skills?
* Success criteria: participant reviews the generated skills, provides feedback, and clarifies any discrepancies

### Post-Test Feedback Questions \[10 min]

* \[Ease of use] How easy was it to interact with Compass and understand its responses?
* \[Accuracy & relevance] How accurate and relevant do you feel the identified skills are? Why?
* \[Accuracy & relevance] Are there any skills that Compass incorrectly identified for you?
* \[Accuracy & relevance] Are there any skills you have that Compass did not identify?
* Did Compass help you gain a clearer understanding of your skills? why or why not?
* Would you recommend Compass to other job seekers? Why or why not?
* In the future, Compass will provide you with advice on which sectors you could focus on for your job searches based on the skills it will have identified for you. What do you think about this?
* In the future, Compass will match you with jobs that align with your identified skills. What do you think about this?
* Any other feedback or suggestions for improving Compass?

Thank you and close


# Compass Customization Guide

Compass supports **white-label customization**, allowing partners to adapt branding, features, language, and data collection so the application looks and feels like their own product.

This guide explains what you can customize and how those changes affect the application.

### Branding

Update the application name, logos, icons, and colors to match your organization’s identity. You can also set search engine details to control how your site appears online.

**Identity**

<div align="left"><figure><img src="/files/OdovDgaXyw8zLAJCMWBr" alt=""><figcaption><p><em>Application identity</em></p></figcaption></figure> <figure><img src="/files/SU9fCa0otKUJf1ycvcEo" alt=""><figcaption></figcaption></figure></div>

**Assets**

<div align="left"><figure><img src="/files/CvyHXJVmtdsykUlSacsz" alt=""><figcaption><p><em>Logo and icon assets</em></p></figcaption></figure> <figure><img src="/files/mojQm3NOUB2Kf3vnZDKQ" alt=""><figcaption></figcaption></figure> <figure><img src="/files/yOipyMCDdhBBtVEDQwWk" alt=""><figcaption></figcaption></figure></div>

**Colors**

<div align="left"><figure><img src="/files/fCToOdyHolg1E0CHdzOu" alt="" width="270"><figcaption><p><em>Branding colors</em></p></figcaption></figure></div>

### Authentication <a href="#authentication" id="authentication"></a>

Decide how users log in. You can configure the authentication process by enabling or disabling login or registration codes.

<div align="left"><figure><img src="/files/XO3reSa4Jprt1Fxksw9A" alt=""><figcaption><p><em>Login and registration with code enabled</em></p></figcaption></figure> <figure><img src="/files/7UlrjwKoWfU7dHvIgSp4" alt=""><figcaption></figcaption></figure></div>

### CV Features <a href="#cv-features" id="cv-features"></a>

Turn the CV functionality on or off. If disabled, all CV‑related options disappear from the application.

<div align="left"><figure><img src="/files/uZKKDErTPbeMU0LwwpFI" alt=""><figcaption><p><em>CV feature enabled</em></p></figcaption></figure></div>

### Skills Report <a href="#skills-report" id="skills-report"></a>

Add your logo(s) to the report, choose the available download formats (PDF or DOCX), and decide which sections are included in the report, including the summary and experience details.

<div align="left"><figure><img src="/files/iyTaPThgyVw3msNpInLL" alt=""><figcaption><p><em>Skills report customization</em> </p></figcaption></figure></div>

### Languages

Select the default language for the application and choose which additional languages users can switch to.

<div align="left"><figure><img src="/files/Dgm8SNoZo3jOPqgm75PV" alt=""><figcaption><p><em>Language switcher interface</em></p></figcaption></figure></div>

### Sensitive Data

Define which personal information Compass collects from users. Available fields include:

* Name
* Contact Email
* Gender
* Age
* Education Status
* Main Activity

Each field can be customized to match your organization’s requirements and translated into supported languages.

<div align="left"><figure><img src="/files/3bYOBleA8iwHWEvZYiG7" alt=""><figcaption><p><em>Sensitive data fields</em></p></figcaption></figure></div>

### How It Works <a href="#how-it-works" id="how-it-works"></a>

Customization is controlled through settings provided during deployment. These settings are passed into Compass automatically and applied across the application once the deployment is complete.

### Important Notes <a href="#important-notes" id="important-notes"></a>

* Only the options listed above can be customized. Core workflows and layouts stay the same.
* If a customization setting is missing or entered incorrectly, Compass will ignore it and use the default settings instead.
* Colors must meet accessibility standards so all content remains readable.


# Roadmap

Our vision for Compass includes:

* ***Expanded geographic reach:*** Widespread adoption of Compass for use in emerging markets
* ***Multilingual support:*** Expand language capabilities to serve diverse populations
* ***Voice integration:*** Enable our users to speak to Compass in their local dialects
* ***Portable skills wallet:*** Compass skills exploration outputs activated in an interoperable skills wallet
* ***Enhanced features:*** Develop more advanced career pathing and skills development recommendations
* ***Integration capabilities:*** Create APIs and tools for seamless integration with existing job platforms and career services
* ***Impact measurement:*** Implement a multi-side randomised control trial to estimate Compass impact on career outcomes


# South Africa

In South Africa, Tabiya partners with Harambee, a not-for-profit social enterprise working on youth unemployment and an anchor-partner for SA Youth, the national employment pathway management network.

## Why South Africa?&#x20;

South Africa faces high youth unemployment rates calling for action into improving intermediation between young job-seekers and employers. Despite being a problem shared with many other countries in the region, the South African case is of particular interest because of the diversity of large scale private and semi-governmental undertakings meant to bring answers to these challenges. Harambee and SAYouth were products of these efforts, and the large database of users, network of stake-holders, and expertise hosted by these organizations present unique opportunities to explore and test a breadth of technical solutions to labour market intermediation.&#x20;

## Our Partner: Harambee Youth Employment Accelerator&#x20;

<div data-full-width="false"><figure><img src="/files/1kxgSWSBIZvQkymNQlY5" alt=""><figcaption></figcaption></figure></div>

Founded in 2011, Harambee Youth Employment Accelerator is a not-for-profit social enterprise working to find solutions to youth unemployment in South Africa, and having expanded operations to Rwanda as of 2018. The organization is an anchor partner for SAYouth, South Africa's national network for youth employment pathway management. Within SAYouth, Harambee operates the multi-channel, tech-enabled SA Youth Platform (sayouth. mobi), which connects a network of 3.8 million jobseeking youth with over 1,300 private and public sector employer partners to facilitate earning and learning opportunities.&#x20;


# Context: the South African Labour Market

Characterized by the world's highest unemployment rates and persistent effects of historical social segregation, the South African labor market presents a particularly challenging context.

{% hint style="info" %}
This review is a work in progress. We will update it continuously with the latest research to date.&#x20;
{% endhint %}

**At 33% as of 2024 Q1, South Africa has the highest unemployment rate in the world, with the rates doubling (61%) for youth aged 15-24.** This is an unequal crisis on multiple dimensions. Rates are overwhelmingly higher among the Black African population (36%), followed a far second colored population (23%). Women face a 13% higher likelihood of unemployment than men.

{% embed url="<https://ourworldindata.org/grapher/unemployment-rate>" %}
*Depending on measures, years and institutions, South Africa continuously displays one of the highest unemployment rate in the world.*
{% endembed %}

**Over three million young people in South Africa are not in employment, education or training (NEET).** The same patterns of unequal outcomes persist: [young South Africans who are not in employment, education or training comprise of a majority black (88%) urban (59%) and women (54%).   ](#user-content-fn-1)[^1]These individual who are NEET are particularly vulnerable to [long spells of unemployment and  insecure work because they do not possess traditional qualifications and experiences that are observed by firms to determine skills and experiences.](#user-content-fn-2)[^2]

{% embed url="<https://ourworldindata.org/grapher/youth-not-in-education-employment-training>" %}
*This map suggests that, contrary to other countries on the continent, the issue South Africa has mostly pertain to labour demand and intermediation  and not education or labour market participation.*&#x20;
{% endembed %}

[**More than 80% of the youth who are NEET have never been employed and have been searching for employment for more than 1 year while 20% have been searching for** ](#user-content-fn-1)[^1]**5+ years.** Low enrolment rates in tertiary education (25.2% as of 2021), is a key factor behind this widespread and persistent unemployability of youth. The lack of tertiary education is further compounded by the apartheid legacy of legal discrimination and segregation that generated large wealth and spatial inequalities across the country. This means, for many of these already disadvantaged jobseekers, available jobs are also located prohibitively far away.&#x20;

{% embed url="<https://ourworldindata.org/grapher/gross-enrollment-ratio-in-tertiary-education?tab=map>" %}

**Together, these factors characterize a labor market facing a vast disconnect between labor demand and labor supply sides.** On the supply side, a large pool of jobseekers struggle to signal their skills to prospective employers and face inherently high search costs due to socioeconomic and spatial inequalities. On the demand side, firms appear uncertain about the quality of the applicants they attract with [many reporting that the skills they are interested in are not emphasized or tested within the education system](#user-content-fn-3)[^3]. This in turn leads to firms relying on trial-and-error recruiting tendencies, creating insecure entry level jobs. [Experimental research found firms more willing to hire when given legal advice on how to fire entry-level workers. ](#user-content-fn-4)[^4]

[**Inaccurate beliefs about skills lead to high prevalence of misdirected job search**](#user-content-fn-5)[^5]**. Job search support interventions have proven useful in improving employment outcomes.** Studies have tested and found positive results for job search support through 1) behavioral encouragement interventions such as [planning support](#user-content-fn-6)[^6], social support; 2) better utilization of job search platforms such as [training jobseekers in how to use LinkedIn](#user-content-fn-7)[^7] and 3) endorsement mechanisms such as [provision of reference letters](#user-content-fn-8)[^8] and[ low cost skill testing and certification to job seekers](#user-content-fn-3)[^3]. When implemented at scale, better job search infrastructure has the potential to bridge the gap between what firms need versus what jobseekers can offer through increased visibility of the labour market and pathways within it for both sides. &#x20;

**Spatial mismatches between jobs and jobseekers, combined with high search costs, further deepen inaccurate beliefs about the job market among youth.** Overly optimistic beliefs about the job market lead to young jobseekers under searching and holding out for "better jobs". On one hand, [ transport subsidies and cash transfers have been found to increase job search activity](#user-content-fn-9)[^9], but these search efforts do not necessarily translate to better employment outcomes on average. [One study found a negative information shock to such subsidies leading to lower reservation wages and acceptance of lower paying jobs.](#user-content-fn-10)[^10]&#x20;

##

## Further readings:&#x20;

Abebe, G., Caria, A.S., Fafchamps, M., Falco, P., Franklin, S. and Quinn, S., 2021. Anonymity or distance? Job search and labour market exclusion in a growing African city. The Review of Economic Studies, 88(3), pp.1279-1310.

Abebe, G., Caria, S.A., Fafchamps, M., Falco, P., Franklin, S., Quinn, S. and Shilpi, F.J., 2023. Matching frictions and distorted beliefs: Evidence from a job fair experiment (No. 958). working paper.

Abel, M., Burger, R., Carranza, E. and Piraino, P., 2019. Bridging the intention-behavior gap? The effect of plan-making prompts on job search and employment. American Economic Journal: Applied Economics, 11(2), pp.284-301.

Abel, M., Burger, R. and Piraino, P., 2020. The value of reference letters: Experimental Evidence from South Africa. American Economic Journal: Applied Economics, 12(3), pp.40-71.

Afridi, F., Dhillon, A., Roy, S. and Sangwan, N., 2023. Social Networks, Gender Norms and Labor Supply: Experimental Evidence Using a Job Search Platform (No. 677). Competitive Advantage in the Global Economy (CAGE).

Banerjee, A. and Sequeira, S., 2023. Learning by searching: Spatial mismatches and imperfect information in Southern labor markets. Journal of Development Economics, 164, p.103111.

Beam, E.A. and Quimbo, S., 2023. The Impact of Short-Term Employment for Low-Income Youth: Experimental Evidence from the Philippines. Review of Economics and Statistics, 105(6), pp.1379-1393.

Bertrand, M. and Crépon, B., 2021. Teaching labor laws: Evidence from a randomized control trial in South Africa. American Economic Journal: Applied Economics, 13(4), pp.125-149.

Bhorat, H., Köhler, T. and de Villiers, D. (2023). Can Cash Transfers to the Unemployed Support Economic Activity? Evidence from South Africa. Development Policy Research Unit Working Paper 202301. DPRU, University of Cape Town.

Branson, N., DeLannoy, A., & Kahn, A. (2019). Exploring the transitions and well-being of young people who leave school before completing secondary education in South Africa. Working Paper Series Number 244, NIDS Discussion Paper 2019/11 Version 1.

Carranza, E., Garlick, R., Orkin, K. and Rankin, N., 2022. Job search and hiring with limited information about workseekers’ skills. American Economic Review, 112(11), pp.3547-3583.

Chakravorty, B., Bhatiya, A.Y., Imbert, C., Lohnert, M., Panda, P. and Rathelot, R., 2023. Impact of the COVID-19 crisis on India’s rural youth: Evidence from a panel survey and an experiment. World Development, 168, p.106242.

Fields, G.S., 2011. Labor market analysis for developing countries. Labour economics, 18, pp.S16-S22.

Franklin, S., 2015. Location, search costs and youth unemployment: A randomized trial of transport subsidies in Ethiopia.

Hardy, M. and McCasland, J., 2023. Are small firms labor constrained? experimental evidence from ghana. American Economic Journal: Applied Economics, 15(2), pp.253-284.

ILO., 202. The impact of COVID-19 on the informal economy in Africa and the related policy responses.

ILO., 2023. African youth face pressing challenges in the transition from youth to work.

Jones, S. and Sen, K., 2022. Labour market effects of digital matching platforms: Experimental evidence from sub-Saharan Africa.

Kiss, A., Garlick, R., Orkin, K. and Hensel, L., 2023. Jobseekers’ beliefs about comparative advantage and (mis) directed search. Available at SSRN 4593303.

Loiacono, F. and Silva Vargas, M., 2023. Improving access to labor markets for refugees: Evidence from Uganda.

McKenzie, D., 2017. How effective are active labor market policies in developing countries? a critical review of recent evidence. The World Bank Research Observer, 32(2), pp.127-154.

Mudiriza, G., De Lannoy, A. (2023). Profile of young people not in employment, education or training (NEET) aged 15-24 years in South Africa: an annual update. Cape Town: Southern Africa Labour and Development Research Unit, University of Cape Town. (SALDRU Working Paper Number 298).

Wheeler, L., Garlick, R., Johnson, E., Shaw, P. and Gargano, M., 2022. LinkedIn (to) job opportunities: Experimental evidence from job readiness training. American Economic Journal: Applied Economics, 14(2), pp.101-125.

World Bank (2023), “Skills and Workforce Development,” worldbank.org.\
Hensel, L., [Tekleselassie](https://cssh.northeastern.edu/faculty/tsegay-tekleselassie/), T., [Isphording](https://sites.google.com/view/ingoeisphording/about-me), I.,  [Radbruch](https://sites.google.com/site/jonasradbruch01/), J. & [Witte](http://www.marcwitte.com/home), M. 2024. Demand for Feedback and Job Search. Working Paper.<br>

[^1]: Mudiriza, G., De Lannoy, A. (2023). Profile of young people not in employment, education or training (NEET) aged 15-24 years in South Africa: an annual update. Cape Town: Southern Africa Labour and Development Research Unit, University of Cape Town. (SALDRU Working Paper Number 298).

[^2]: Branson, N., DeLannoy, A., & Kahn, A. (2019). Exploring the transitions and well-being of young people who leave school before completing secondary education in South Africa. Working Paper Series Number 244, NIDS Discussion Paper 2019/11 Version 1.

[^3]: Carranza, E., Garlick, R., Orkin, K. and Rankin, N., 2022. Job search and hiring with limited information about workseekers’ skills. American Economic Review, 112(11), pp.3547-3583.

[^4]: Bertrand, M. and Crépon, B., 2021. Teaching labor laws: Evidence from a randomized control trial in South Africa. American Economic Journal: Applied Economics, 13(4), pp.125-149.

[^5]: Kiss, A., Garlick, R., Orkin, K. and Hensel, L., 2023. Jobseekers’ beliefs about comparative advantage and (mis) directed search. Available at SSRN 4593303.

[^6]: Abel, M., Burger, R., Carranza, E. and Piraino, P., 2019. Bridging the intention-behavior gap? The effect of plan-making prompts on job search and employment. American Economic Journal: Applied Economics, 11(2), pp.284-301.

[^7]: Wheeler, L., Garlick, R., Johnson, E., Shaw, P. and Gargano, M., 2022. LinkedIn (to) job opportunities: Experimental evidence from job readiness training. American Economic Journal: Applied Economics, 14(2), pp.101-125.

[^8]: Abel, M., Burger, R. and Piraino, P., 2020. The value of reference letters: Experimental Evidence from South Africa. American Economic Journal: Applied Economics, 12(3), pp.40-71.

[^9]: Bhorat, H., Köhler, T. and de Villiers, D. (2023). Can Cash Transfers to the Unemployed Support Economic Activity? Evidence from South Africa. Development Policy Research Unit Working Paper 202301. DPRU, University of Cape Town.

[^10]: Banerjee, A. and Sequeira, S., 2023. Learning by searching: Spatial mismatches and imperfect information in Southern labor markets. Journal of Development Economics, 164, p.103111.


# Employment Journeys

The Employment Journey (EJ) dataset compiled by Harambee is a valuable source of information about job-seekers' previous experience.

## &#x20;What is the EJ data?

Harambee, through SAYouth, regularly collects updated employment information from youth enrolled on their platform. This information, collected over XX years, produces a rich database of [activities and occupations](#user-content-fn-1)[^1] held  by enrolled youth across time, highlighting the diversity of **employment journeys (EJ)** among young South African job seekers.  For each activity and occupation held by young job-seekers, the database contains titles, a brief job description and additional information on conditions of employment/engagement such as contract duration and permanence, business classification (product/service/both) etc.&#x20;

## What do employment journeys teach us?&#x20;

### <mark style="color:blue;">Employment experiences of young South African job seekers are a mix of formal, informal, and</mark> [<mark style="color:blue;">unseen</mark>](https://docs.tabiya.org/) <mark style="color:blue;">experiences.</mark>&#x20;

We first look at the distribution of different types of livelihoods in the EJ data. As a first, cut, the data is decoded into three subsets for those who report any occupation or activity:&#x20;

1. Those who report working for someone else, i.e. as being “employed” or having an “employer” are considered as a sample of the formal employment sector.
2. Those who report working for themselves i.e. "self-employed" are classified as [microentrepreneurs](/tabiya-south-africa/seen-economy/entrepreneurial-skills) and capture a sample of individuals from South Africa’s informal economy.&#x20;
3. Volunteer work represents unpaid activities reported by jobseekers, and this includes a range of individual or community level unpaid work starting from apprenticeships to community service.

The formal and informal sectors together represent what we are calling the [seen economy](/tabiya-south-africa/seen-economy), while volunteer work and the 57.7% unemployed and unspecified sample represent the [unseen economy](/tabiya-south-africa/unseen-economy). Strikingly, among those who have been involved in some work in the past 30 days, **informal work and volunteer work make up 60% of cases**. This confirms that, in the South African case, recognizing and highlighting skills acquired by young job-seekers in the unseen and informal economies is crucial for inclusivity.

**Table 1: Distribution of Harambee EJ data across livelihood types**

|                                   |                  |                  |                                                       |
| --------------------------------- | ---------------- | ---------------- | ----------------------------------------------------- |
| Livelihood type                   | Observations (N) | % of full sample | [% of employed viable sample](#user-content-fn-2)[^2] |
| Formal sector                     | 199,399          | 16.8%            | 39.2%                                                 |
| Microentrepreneur/informal sector | 161,392          | 13.6%            | 31.7%                                                 |
| Volunteer work                    | 147,981          | 12.5%            | 29.1%                                                 |
| Unemployed                        | 664,663          | 56.0%            | -                                                     |
| Unspecified                       | 13,853           | 1.2%             | -                                                     |
| Total                             | 1,187,288        | 100.0%           | 100.0%                                                |

*Source: Harambee EJ Data*

<figure><img src="/files/jcZN0p99xxaZPjgT2XXm" alt=""><figcaption><p>Figure 1: <strong>Distribution of Harambee EJ data across livelihood types;</strong> Source: Harambee EJ Data</p></figcaption></figure>

{% hint style="info" %}
**Note on data cleaning:** As evident from Figure/Table 1, majority of individuals in the Harambee EJ data are classified as “unemployed”. Although individuals in this category do still occasionally report a job title, this is non-sensible. As such, we remove all unemployed individuals from our analysis and further remove those individuals whose livelihood type was categorized as “Unspecified”. This leaves a total of approximately 508 000 individual EJ records to analyze, of which 39.2% are in the formal sector, 31.7% in the informal/microenterprise sector, and 29.1% in volunteer work.  &#x20;
{% endhint %}

### <mark style="color:blue;">Noting data limitations and sample comparability with the SA QLFS</mark>

It is important to note that the EJ data, while a rich source of information, is available for a very specific selection of individuals: South African youth who have registered as jobseekers on the Harambee platform, with a disproportionate number of individuals located in and around the province of Gauteng in South Africa (see Table 2). Having an overrepresented urban sample could introduce biases in our analysis, as [our assessment of whether ESCO is applicable to our context](/tabiya-south-africa/seen-economy/alternative-titles) will be based on livelihood descriptions provided by these individuals.

To avoid such biases, we are making sure that the [Harambee platform (version 0.1) ](/tabiya-south-africa/harambees-platform-version-0.1)**allows users to use free text to describe their experiences**. Thanks to this, all livelihoods types can be taken into account, and the new information provided can enrich Harambee's list of activities and occupations.

**Table 2: Provincial breakdown of youth from Harambee EJ data relative to South African Quarterly Labour Force Survey data**

<table data-header-hidden><thead><tr><th></th><th width="153"></th><th></th><th></th><th></th></tr></thead><tbody><tr><td></td><td>Harambee EJ Data</td><td>Harambee EJ Data</td><td>QLFS 2023Q1</td><td>QLFS 2023Q1</td></tr><tr><td></td><td>Number</td><td>Percent</td><td>Number</td><td>Percent</td></tr><tr><td>Eastern Cape</td><td>102,750</td><td>8.94</td><td>2,019,113</td><td>11.55</td></tr><tr><td>Free State</td><td>33,826</td><td>2.94</td><td>823,531</td><td>4.71</td></tr><tr><td>Gauteng</td><td>423,689</td><td>36.88</td><td>4,483,754</td><td>25.65</td></tr><tr><td>KwaZulu-Natal</td><td>219,352</td><td>19.09</td><td>3,459,500</td><td>19.79</td></tr><tr><td>Limpopo</td><td>91,047</td><td>7.93</td><td>1,821,289</td><td>10.42</td></tr><tr><td>Mpumalanga</td><td>94,574</td><td>8.23</td><td>1,395,100</td><td>7.98</td></tr><tr><td>North West</td><td>54,484</td><td>4.74</td><td>1,153,158</td><td>6.60</td></tr><tr><td>Northern Cape</td><td>25,790</td><td>2.25</td><td>361,136</td><td>2.07</td></tr><tr><td>Western Cape</td><td>97,518</td><td>8.49</td><td>1,965,682</td><td>11.24</td></tr><tr><td>Unspecified</td><td>5,728</td><td>0.50</td><td>-</td><td></td></tr><tr><td>Total</td><td>1,148,758</td><td>100</td><td>17,482,262</td><td>100</td></tr></tbody></table>

Source: Own calculations using Harambee EJ Data and Statistics South Africa (2023)

&#x20;*Note: 1. The National Youth Commission (1996) defines youth as all individuals between the ages of 15 and 34 (inclusive). Due to the age of majority in South Africa being 18, we restrict our analysis to youth aged between 18 and 34 (inclusive). 2. Numbers from QLFS weighted using sampling weights.*

### <mark style="color:blue;">Analyzing job transitions shed light on how moves are influenced by past experiences; impact of job moves on economic stability, mobility and fragility; and employer preferences.</mark>&#x20;

**About the data and analysis:** The job transition analysis is based on SAYouth placement data of 164,527 unique youth who report multiple opportunities on the platform between 2012 through May 2024. These youth make up 4.3% of the 3.8 million youth who are registered on the platform, and 18.6% of the 884,000 unique youth that report at least [one opportunity on the platform](#user-content-fn-3)[^3] (with 1.2 million total opportunities).&#x20;

The analysis examines **the first transitions** made by these youth. This means, opportunities are only included if they have a valid start date and a valid opportunity type namely Public Employment Program (PEP), Formal Sector (FS), or Make Your Own Money (MYOM). Among youth who report multiple opportunities, 80% record two so we capture their full journey in this sample. For the remaining 20% with more than two reported opportunities, we examine only the first two.&#x20;

{% hint style="info" %}
Note that this analysis does not account for:&#x20;

**Youth who have a first opportunity and then do not transition to something else.** When such gaps are present on a user’s profile, we cannot conclude whether the user became unemployed or simply did not update their SA Youth profile. For this reason, we do not consider them in this analysis.

**Youth who gained employment and retained it, without transitioning to anything else.** Youth who secure a job and stay in it for many months and years are positive outliers among South Africa’s youth. These youth are not studied in this analysis, however, because there is no transition to report.
{% endhint %}

A total of 329,054 placements in the data have been classified into one of three types of opportunities: PEP, MYOM/microenterprise, and formal sector. Of the three, the sample is PEP-heavy.&#x20;

The analysis examines opportunities starting as early as 2012, but the majority fall within the 2020-2022 period. For 27% of transitions, the time between start dates of 1st & 2nd opportunities is less than 100 days. For 71% of transitions, this time between is less than 400 days.&#x20;

<figure><img src="/files/tOMb5zRfRuYomz1oVWVc" alt=""><figcaption><p><strong>Figure 2: Distribution of opportunities across employment categories; Source: Harambee EJ Data</strong></p></figcaption></figure>

#### *<mark style="color:green;background-color:yellow;">Learnings on flows: Youth zigzag across many types of opportunities, but the likelihoods of different transitions are influenced by and in most cases, mirror their previous experiences.</mark>*&#x20;

**Transitions overall, across PEP, FS and MYOM, are more likely to be within groups (60.3%) than between groups.** PEP participants tend to transition to another PEP 71% of the time, and youth in the formal sector tend to transition to another role in the formal sector 54% of the time. Meanwhile, youth in MYOM opportunities are much less likely to transition to another MYOM opportunity (23%) than to the formal sector or PEPs.

<figure><img src="/files/gAXsrmmJDZgD3JLeMwfJ" alt="" width="563"><figcaption><p><strong>Figure 3: Flow diagram of transitions within and across groups; Source: Harambee EJ Data</strong></p></figcaption></figure>

* **Transitions from/to PEP**: PEPs are the most common opportunities, with DBE as the single largest opportunity. 70% of youth in this sample participate in a PEP as a first and/or second opportunity. On net, there is transition out of PEP. However, 54% of all youth in the sample transitioned to a PEP, composed of 71% youth with prior PEP experience, 36% from MYOM and 29% from the formal sector.&#x20;
* **Transitions from/to formal sector**: The formal sector is the second most common type of opportunity overall, with 17% of all youth in this sample staying in the formal sector across both opportunities. On net, there is more transition into the formal sector. 32% of all youth in the sample transitioned to the formal sector, composed of 71% youth with prior FS experience, 41% from MYOM and 18% from PEP.&#x20;
* **Transitions from/to formal sector**: On net, more youth move into MYOM than moving out of. However, importantly, youth who were engaged in MYOM were nearly twice as likely to transition to the formal sector or PEPs than to a second MYOM opportunity. Of the 14% that transitioned into MYOM, 23% were previously engaged in MYOM, 17% from FS and 11% from PEP.&#x20;

[^1]: "activities" refer to volunteer work, "occupations" refer to paid work

[^2]: \*excluding unemployed/ unspecified

[^3]: Opportunities in this dataset are self-reported, partner-reported, or reported by Harambee. “Ecosystem” opportunities reported in bulk by partners like DBE or GBS are not included in this data.&#x20;


# Time-Use

Understanding how young jobseekers spend their time highlights the main activities through which they might develop valuable human capital.

Time-use surveys worldwide aim to capture how individuals spend their day. The International Classification of Activities for Time Uses Statistics (ICATUS) is an international standard that provides a framework and classification for these time-use activities. The taxonomy encompasses a comprehensive set of activities,  including those considered typically productive such as doing a job or working for pay, as well as unproductive activities like sleeping or watching television.&#x20;

The ICATUS taxonomy offers a more inclusive approach to defining what is productive. It encompasses activities that fall outside the production boundary set by the [System of National Accounts](https://unstats.un.org/unsd/nationalaccount/sna.asp). These include services rendered without pay for own and others' final use, such as cleaning one’s own dwellings or cooking one’s own food. These activities are considered productive insofar that they have paid equivalents (for instance, a professional cleaner or a cook). Therefore, they engage human capital and enable human capital formation.&#x20;

We identify four types of such activities in ICATUS, which we call “[unseen](/tabiya-south-africa/unseen-economy)” activities.

1. **Unpaid domestic services for household and family members (Div. 3)**
2. **Unpaid caregiving services for household and family members (Div. 4)**
3. **Unpaid direct volunteering for other households (Div. 51)**
4. **Unpaid community- and organization-based volunteering (Div. 52)**

<figure><img src="/files/QVzeZwafZSsbVeBeHLui" alt=""><figcaption><p>Figure 1:  Time Use of Youth (15-34 years old) in South Africa, Excluding Sleep</p></figcaption></figure>

Highlighting this human capital in the South African first entails uncovering through which activities it might be formed. The South African time use survey from 2011 is the latest national wide source of data on how South Africans spend their time on the daily. When it comes to unseen activities, caring for someone in one’s household and maintaining one’s household takes up 19,9% of young South African’s time (15-34 years old). On the other hand, community service for non-household members such as volunteering (at Church, for instance) takes up 0,4% of young South Africans’ time.


# Seen Economy

"Seen" activities are those which are accounted for as productive (including informal work) in the System of National Accounts (SNA). \[to be edited later]

When adapting ESCO, a European labor taxonomy, to the South African "seen" economy, the first step  involved asking if occupations in Europe - and the skills associated with them - are similar to those in South Africa. The answers to this were approached from two directions. &#x20;

First, do occupations with the same titles across Europe and South Africa substantively entail the same tasks and require the same skills? For instance, being a "specialize seller" or "taxi driver" in Europe (and hence in ESCO), rarely involves having to negotiate prices or bargain. However,  these are critical skills for such occupations in most developing countries, including South Africa. These gaps in occupational definitions and skills requirements are larger for more informal occupations, especially in a gig economy like South Africa's.&#x20;

Second, do Europeans and South-Africans refer to the same occupations in the same terms? For example, is someone who conducts data collection called an data collector in both regions? Locals in South Africa use the term "fieldworker" more commonly. ESCO provides space to link "alternative titles" to every occupation it lists, so as to reflect this diversity of names given to the same jobs. It is also translated in every European language and reflects the conceptual diversity of ways languages refers to the same occupations. Expansion of ESCO beyond the EU, particularly in relatively data-poor developing contexts, requires careful groundwork for the new countries and languages.&#x20;

Tabiya approached answers to these questions in two ways:&#x20;

1. Using existing databases and taxonomy: [We used a combination of Harambee's rich database of employment data, and SA's pre-existing local labor taxonomy in testing methodologies to systematically add South African alternative titles to ESCO.](/tabiya-south-africa/seen-economy/alternative-titles)
2. Directly ground truthing the applicability and usability of ESCO for microentrepreneurs: [We conducted a survey with young micro-entrepreneurs registered with Harambee](/tabiya-south-africa/seen-economy/entrepreneurial-skills) in trying to assess whether: i) respondents could readily find their work on a list of ESCO occupations ii) skills under selected ESCO occupations fully captured the skills that a microentrepreneur thinks they have/need. &#x20;


# Alternative Titles

Including all diverse experiences and skills held by young South African entails ensuring they identify with the titles of occupations provided on our platform.

Alternative titles for occupations allow us to make sure that **there is no "one best way" to refer to a job, an informal activity, or a hobby**. To make sure every one feels represented on the Harambee platform, we included idiomatic occupation titles, and allow for an **adaptative listing of alternative occupation titles**.&#x20;

## Why are alternative job titles important?&#x20;

In any economy, individuals who are otherwise employed to do the same job may be **hired under different titles**. The ESCO taxonomy attempts to address this concern by providing a list of **alternative titles** that are associated with the different occupations captured in its database. However, regional variations in job titles may exist, and so while the ESCO list of so-called “alternative titles” or “alternative labels” for occupations may be sufficient for the European labour market, the same cannot necessarily be said for the South African labour market.&#x20;

{% hint style="info" %}
The ESCO taxonomy exists for **all European languages**: the same occupations and skills thus have different names. However, problems arise from the fact that the occupations and skills may be called differently depending on the **regional variations of the English language**. For instance, in South-Africa, a “survey enumerator” may be called “fieldworker”. Therefore, one needs to identify, manually or thanks to algorithmic text processing, alternative titles that are relevant in each regional context.&#x20;
{% endhint %}

## Localizing ESCO to the South African Labor Market

In localizing ESCO to South Africa's context, we aimed to identify to what extent ESCO's existing titles and alternative titles captured the jobs done by participants in this labor market. The exploration and subsequent alternative title additions were completed in two blocks of work: 1) a matching exercise using the EJ data entries and ESCO titles 2) a cross-referencing exercise using SA's local labor taxonomy, the Organizing Framework for Occupations (OFO), with ESCO.&#x20;

### 1. Matching with EJ Data&#x20;

To get a sense of the need to add alternative titles for ESCO occupations, we first mapped 500 formal, seen occupations from the[ EJ data](/tabiya-south-africa/context-the-south-african-labour-market/employment-journeys) to the ESCO taxonomy. The first attempt was an automated text-matching exercise between job titles and descriptions in the EJ data and the ESCO taxonomy. Given the diversity of job titles and job descriptions on both sides, automating the text-matching proved to be a significant challenge. Consequently, a manual review of Harambee's EJ data needed to be undertaken as an initial proof of concept before large-scale data collection and analysis could be undertaken.&#x20;

Random samples of EJ entries from both formal sector workers and the microentrepreneurs were analyzed. A preliminary analysis of a random sample of 50 EJ entries from each category was followed by a scaled up analysis of [500 entries](#user-content-fn-1)[^1] from each category.  Results were consistent across both subsamples. For brevity, the results presented in the following section are based on the random subsample of size 500.

To match the job descriptions provided to by young job-seekers to ESCO occupations, we manually applied the following rules:&#x20;

{% hint style="success" %}
**Rule 1:** If there was an exact match in the user-provided livelihood title from the EJ data, and either an ESCO occupation title, or an ESCO occupation’s alternative title, then the individual was matched to this ESCO occupation code at the 4.1-digit level.&#x20;
{% endhint %}

{% hint style="success" %}
**Rule 2:** Where user-provided livelihood titles were etymologically linked to an ESCO title or alternative title by a common “root word” – e.g., the user-provided title of “Teacher assistant” is etymologically linked to the ESCO title “Teaching assistant” through the common root word “teach” – then these individuals were classified as directly mappable to ESCO without the need for an additional alternative title.
{% endhint %}

{% hint style="success" %}
**Rule 3:** Where there was no directly mappable link between user-provided livelihood titles and ESCO occupation titles, individuals were only allocated to a given occupation code if we could be reasonably sure that we had identified the job the individual was trying to communicate through their EJ livelihood data.
{% endhint %}

{% hint style="success" %}
**Rule 4:** Where there was no direct link between user-provided livelihood title and ESCO occupation title, and it was not clear what occupation was being described by the user-provided livelihood title, we opted to not map occupations.&#x20;
{% endhint %}

~~This manual mapping resulted in a total of 279 individual EJ entries being matched to 73 unique ESCO occupation codes. At the occupation level, 51 occupations required no additional alternative titles to be added, with 39 providing exact matches to individuals’ user-provided titles, and 12 providing root-word matches to user-provided titles. The remaining 22 occupations potentially require the addition of alternative titles to ensure cohesiveness between South African occupation titles and ESCO occupation titles.~~&#x20;

~~At the individual-level, we find that 224 individual user-provided occupation titles were matched to ESCO occupation titles or alternative titles via either an exact match or root word match. This accounts for 80.3% of the total matched sample of individuals. In other words, only 19.7% of individuals in the matched sample entered occupation titles that may require the addition of alternative titles to the ESCO database. This information is summarized in Table 5.~~

The exercise resulted in a 80.3% match (exact + root-word matches) at the individual level and a 70% match at the occupation level. Table 5 summarizes the findings.    &#x20;

**Table 5: Summary of user occupation title match to ESCO occupation title, by occupation group.**

<table data-header-hidden><thead><tr><th width="319"></th><th width="198"></th><th width="174"></th><th width="110"></th><th width="169"></th><th></th></tr></thead><tbody><tr><td></td><td>Number of individuals employed in occupation category</td><td>Number of individuals matched exactly</td><td>Individuals as % of all matched individuals</td><td>Number of occupations</td><td>Occupations as % of all matched occupations</td></tr><tr><td>Occupation matches exactly to ESCO title or alternative title</td><td>94</td><td>93</td><td>33.3%</td><td>39</td><td>53.4%</td></tr><tr><td>Occupation has a root word match to ESCO title or alternative titles</td><td>107</td><td>88</td><td>31.5%</td><td>12</td><td>16.4%</td></tr><tr><td>Occupation likely requires addition of alternative title</td><td>78</td><td>43</td><td>15.4%</td><td>22</td><td>30.1%</td></tr><tr><td>Total</td><td>279</td><td>224</td><td>80.3%</td><td>73</td><td>100.0%</td></tr></tbody></table>

*Source: Harambee EJ data and own calculation*

{% hint style="warning" %}
Note: Although there are 94 individuals who are in occupations with exact matches to ESCO titles or alternative titles, only 93 individuals had exact matches to ESCO titles. This is because one individual described their job as “Youth program – cleaning and helping them with their homework”. This individual has been allocated to occupation “Child care worker”. While this is not an exact match for the individual, their user-provided occupation title is a description rather than a title.&#x20;
{% endhint %}

## The South African Organizing Framework for Occupations

[In our EJ individuals sample, we find a total of 55 people who do not have exact matches to an ESCO occupation title or alternative title](#user-content-fn-2)[^2]. In a further attempt to match these occupations, we opted to explore using the South African Organizing Framework for Occupations (OFO) as a cross-referencing tool for potential alternative titles that would need to be added to the ESCO framework.&#x20;

<details>

<summary>Why we chose to use the OFO instead of the South African Standard Classification of Occupations (SASCO) framework? </summary>

As an alternative to the OFO, the other local labor taxonomy available to the team was the South African Standard Classification of Occupations (SASCO) framework, which is based on ISCO and provides information on tasks related to occupations.&#x20;

The OFO was decidedly better suited for the purpose of this exercise than the SASCO for several reasons:&#x20;

1. SASCO was last updated in 2012, while the OFO is more regularly updated, the latest being 2021. Time-relevant frameworks are particularly important in dynamic and gig-economy driven labor markets like that in SA.&#x20;
2. Unlike the OFO, SASCO does not readily provide a list of alternative titles. In trying to develop an exhaustive list of different titles for the same occupation, the team opted for the framework with a larger existing base of alternative titles.&#x20;
3. Although the national statistics agencies uses the SASCO classification to generate labour market statistics, the South African Department of Higher Education and Training (DHET) notes that SASCO does not provide the level of detail required for skills-related planning and interventions (DHET, 2013). As a result, DHET developed the OFO framework to monitor skills supply and demand in South Africa, which aligns more closely with Tabiya and Harambee's objectives as well.&#x20;

</details>

**Table 6: List of ESCO occupations with suggested additional alternative titles (based on sample of 500)**

| ESCO code | ESCO title                       | Alternative title 1  | Alternative title 2 |
| --------- | -------------------------------- | -------------------- | ------------------- |
| 2146.5    | Metallurgist                     | Aluminium maker      |                     |
| 2330.1    | Secondary school teacher         | Substitute educator  |                     |
| 2341.1    | Primary school teacher           | Substitute educator  |                     |
| 2342.1    | Early years teacher              | Substitute educator  |                     |
| 2342.2    | Freinet school teacher           | Substitute educator  |                     |
| 2342.3    | Montessori school teacher        | Substitute educator  |                     |
| 3139.1    | Automated assembly line operator | Production assembler |                     |
| 3312.1    | Bank Account manager             | Universal banker     |                     |
| 3341.6    | Field survey manager             | Fieldwork supervisor |                     |
| 3343.1    | Administrative assistant         | Office administrator | Chief invigilator   |
| 4110.1    | Office clerk                     | Control center clerk |                     |
| 4212.7    | Odds compiler                    | Fixed odds clerk     |                     |
| 4227.2    | Survey enumerator                | Surveyor             | Fieldworker         |
| 5142.2    | Beauty Salon Attendant           | Beauty advisor       |                     |
| 5223.6    | Shop assistant                   | Store assistant      | Till packer         |
| 5230.1    | Cashier                          | Till operator        |                     |
| 5312.1    | Early years teaching assistant   | Educator assistant   | ECD assistant       |
| 7321.1    | Prepress technician              | Scanning engineer    |                     |
| 7422.7    | Telecommunications technician    | DSTV installer       |                     |
| 9212.4    | Livestock worker                 | Milker               |                     |

*Source: Own construction based on Harambee EJ data and European Commission (2022).*

The above exercises demonstrated that for the most part, ESCO titles and alternative titles are able to capture most of South Africa’s occupations, and the results can be improved by additionally using the OFO framework.&#x20;

We then chose to merge the 2021 OFO list of occupations, specialisations and alternate titles to ESCO. Among individuals in our EJ data sample who provided sufficient information to enact a mapping, 89% of these can find their occupation description in some combination of ESCO and OFO titles. This makes a substantive case to use an OFO-ESCO merge as a baseline from which to build our taxonomy for the South African labour market.

<details>

<summary>Challenges in merging the OFO and ESCO taxonomy</summary>

The OFO and ESCO taxonomies diverge at the 4-digit level. The two frameworks create different categorisations of occupations at the more granular levels. Thus, there is not necessarily a clean one-to-one mapping between [the 6-digit level OFO classification of an occupation, specialisation, or alternative title and the ESCO 4.1-digit occupation classification](#user-content-fn-3)[^3].&#x20;

These discrepancies in how the additional granularity is added meant that the mapping was not a case of simply matching the occupation codes to one another. Therefore, to ensure accurate mapping at these levels, either a string search match, or a manual mapping were required.&#x20;

</details>

The matching from ESCO to OFO was done in two steps. First, straightforward matches were made thanks to a simple matching algorithm the following two rules:&#x20;

{% hint style="success" %}
If the exact equivalent occupation (same title) exists in ESCO and OFO, a match happens.&#x20;
{% endhint %}

{% hint style="success" %}
If the ESCO title of an occupation is included as an alternative title for an occupation in OFO, then a match happens.&#x20;
{% endhint %}

## End product: OFO-ESCO merge

The approach finally adopted was manual matching of the remaining OFO occupations to ESCO, through identification of close conceptual equivalents. This work contributes to two downstream functions: First, it allows our partners to build platforms around a localized taxonomy that speaks to young jobseekers, thus improving the functioning of the website. Second, it enriches functionality of [Compass](broken://spaces/0O0RbZr6qGsjDH7HKDCo), Tabiya's interactive chatbot meant to help young job seekers identify their skills.&#x20;

The key to maintaining a time-relevant labor taxonomy is to build systems which are consistently taking feedback from users. Using a fixed taxonomy runs the risk of future users not finding occupations and skills because of both new kinds jobs being created and existing jobs being rebranded. To manage these risks, Harambee continues to include a free text option for job descriptions/titles on the platform. This option doubles input and feedback for the research team at Tabiya to continue adding missing occupation titles and skills to the taxonomy.&#x20;

[^1]:

[^2]: This includes the individual who inserted a job description instead of a job title when prompted for a title, thus actually making it 54 individuals with functionally different occupation titles that may require the addition of alternative titles.

[^3]: These are the appropriate levels to compare the occupational structure to as they both provide one additional layer of granularity from the 4-digit ISCO classification.&#x20;


# Entrepreneurial Skills

To validate our approach for the informal economy, we conducted a survey among a subset of South African micro-entrepreneurs registered to the Harambee website.

To assess the applicability of the ESCO taxonomy to the informal micro-entrepreneur economy of South Africa, we conducted a series of primary data collection exercises with Harambee users. The core objective of this data collection was to finally address two key questions:

1 - Can all activities in the informal micro-entrepreneur economy in South Africa be associated to at least one ESCO occupation?&#x20;

2 - Does the ESCO taxonomy effectively encompass both occupation-specific skills and entrepreneurial skills prevalent in the informal micro-entrepreneur economy?

{% hint style="success" %}
Throughout the data collection process, we collaborated closely with the Harambee team to ensure its relevance to the youth in South Africa. This collaboration encompassed various stages, including conceptualization, survey instrument design, and sampling.
{% endhint %}

## <mark style="color:blue;">Activity 1: Focused Group Discussions to identify a relevant list of general entrepreneurial skills</mark>&#x20;

Microentrepreneurs constitute a unique segment within the workforce, as they possess two distinct sets of skills. First, they possess occupation-specific skills directly related to their trade; for instance, a hairdresser possesses skills such as hair washing and styling. Second, they exhibit general entrepreneurial skills required for self-employment or own-account work, such as personal initiative and drive, identifying suppliers, resource management, and understanding customer needs.&#x20;

Harambee facilitated focused group discussions (FGD) with **18-35 year-old, prior or current informal micro-enterprise owners** within their network. These FGDs were designed to identify the second set of **generic entrepreneurial skills** that might be relevant to microentrepreneurs in South Africa

[A sample of 60 participants including replacements was selected, balanced on gender, age group and type of entrepreneurial venture reported](#user-content-fn-1)[^1]. Ten percent people with disabilities were included in the sample. To encourage participation, a circular was disseminated, inviting eligible individuals to join the FGDs, with selected participants offered R450 for their involvement

[Six sessions with a total of 33 participants were conducted over three days at the Harambee Braamfontein offices in Johannesburg.](#user-content-fn-2)[^2] Each session featured diverse entrepreneurial activities among participants, ranging from car washing and music production to tutoring, baking, goods selling, and hairdressing. The FGDs were designed as a series of individual exercises combined with group discussions, aiming to gain insights into how young individuals perceive their skills as microentrepreneurs.&#x20;

<details>

<summary><mark style="color:purple;">Deep-dive - designing the sessions: each group discussion was structured to allow elicit individual and group feedback from microentrepreneurs, while also extracting insights on how they perceived these general skills</mark> </summary>

The discussions were structured across 5 activities:

**Activity 1: Job descriptions and job titles.** The objective of this exercise was to elicit how people describe the venture that they run and how they identify themselves in the micro-entrepreneurial space. Participants were asked to write down the different ways they would describe themselves with respect to their business/hustle if they had to put down/look up a job title - one sticky note per title.

**Activity 2: Open Entrepreneurial Skills Elicitation.** Participants are oriented on how running a business or hustle of their own entails practicing some specific job-related skills but also some general entrepreneurial skills. The objective of this activity was to elicit a collection of general entrepreneurial skills that the participants feel they have needed/need to run their own ventures.

**Activity 3: Skills List Resonation.** Participants were handed 3 lists of skills, which they are asked to assess individually first. On each list they were to mark out which skills they possessed, skills they do not possess but feel are relevant, skills they feel are unclear/irrelevant to entrepreneurs. Once the three lists were completed, participants were asked to choose which list they resonated with most.

**Activity 4: Learned vs. Unlearned skills.** This segment aims to get an understanding of which skills are effortful and which ones come more naturally to participants. The exercise is in effect done in two parts: a) group elicitation exercise where the facilitator jots down learned and unlearned skills being called out by participants on the white board + open discussion on the skills being called out; b) as part of activity 5, participants take back the skills they had put up on activity 2 and arrange them on their individual white charts based on their own perception of whether they are learned/unlearned skills.

**Activity 5: Individual reflections.** Organize individual charts given to reflect inputs. This allowed participants to re-interact with and organize all the pieces they had put up throughout the exercises. At the end, participants were asked to share overall reflections on the exercise in terms of what they learned, the importance of entrepreneurial skills etc. &#x20;

After the first day, sequencing of Activity 2 and 3 were interchanged. Therefore, Groups 1, 2, and 3 completed the skills list resonation after being primed on entrepreneurial skills through the skills elicitation exercise, and Groups 4, 5 and 6 completed the same in the opposite order i.e. the skills list resonation was completed without any priming discussion on what entrepreneurial skills can include (details on this decision in the next section).

</details>

Participants received three skill lists from ESCO and EntreComp[^3], indicating skills they had, skills they didn't have but deemed relevant, and irrelevant or unclear skills for entrepreneurs. They then chose the lists, and skills within lists, that resonated most with their work.&#x20;

<details>

<summary><mark style="color:purple;">Creating the skills lists: weighing tradeoffs across needing to combine separate frameworks (ESCO + EntreComp), and attempting to generate a list within ESCO</mark> </summary>

Tabiya's taxonomy development so far, had been fully contained within the ESCO frameworks, involving repurposing and expanding definitions/titles for pre-existing skills within it. However, when it came to generic microentrepreneurial skills, ESCO did not provide a clear pathway. Some occupations which would most frequently align with these jobs, such as "specialized sellers" listed some of the generic skills we needed but this was incomplete since these were formulated to describe specific occupations, as opposed to a category of employment type.&#x20;

This raised two questions:&#x20;

1. Should we use the EntreComp framework for reporting of generic entrepreneurial skills?&#x20;
2. Would we be able to create a list of generic entrepreneurial skills from existing ESCO skills that functions comparably or better than EntreComp?&#x20;

&#x20;For the first, we included the list of 60 threads (competencies) from EntreComp in our exercise as one of the candidates. For the second question, there were two ways we generated the ESCO list: one was to manually survey all skills under most-relevant microentrepreneurial occupations and pick out all generic entrepreneurial skills. The alternate was a text-matching exercise between EntreComp competencies and descriptors with ESCO skills and descriptors may offer a more systematic, algorithmic way of generating the list.&#x20;

The tradeoffs were that the first offered potentially high local contextualization, and relatively straightforward to do, but it was based on subjective and arbitrary assessment. The second, while more systematic on paper, would require significantly more effort to generate a sensible list. Even then, depending on the success of the text matching, may require human intervention.  &#x20;

The exercise tested all three lists.&#x20;

</details>

**The EntreComp list was most preferred, with 16 out of 33 participants (48%) choosing it as the list they resonated with the most.** Participants commonly cited the relatability of the EntreComp list as the primary reason for its preference, noting that they could identify with most of the skills on that list. Other positive feedback highlighted its conciseness, clarity, and self-explanatory nature; emphasis on interpersonal skills related to business; and better inclusion of ideas, creativity, and motivation compared to the other lists.

### Creating the final list

Based on the FGD findings, EntreComp was a clear starting point for entrepreneurial skills. However, the primary objective of the data collection exercise was to assess the relevance of the ESCO framework in this domain. So a hybrid approach was adopted in formulating the final list: i) each skill in the EntreComp list was manually matched to the closest skill in the ESCO framework; ii) this list was then further enriched by incorporating skills from the ‘retail entrepreneur’ ESCO occupation. The final list included ### skills, collectively serving as the list of potential entrepreneurial skills for the primary data collection.

## <mark style="color:blue;">Activity 2: Collecting descriptions of micro-entrepreneurial activities via an online survey</mark> &#x20;

The second step of our work with Harambee was about deepening our insights into income-generating activities undertaken by young South Africans.&#x20;

To do so, we developed a [short online survey](https://www.dropbox.com/scl/fi/jk74kqeqke0ll5pg3dda6/IL_Micro-entrepreneurs_Survey1_Questionnaire_v1.0.pdf?rlkey=30qr4yi206vmttyyjx0uroh4n\&dl=0) for individuals within the SA Youth network. The primary aim of the survey was to obtain accurate descriptions of the main micro-entrepreneurship activities individuals engaged in to earn income in the past 30 days. To achieve this, we included the following question:&#x20;

> ***"Tell us about the main way you make money. What things or services do you sell? Write a short sentence or two explaining what you do"***

We also collected current contact details, the number of income-generating activities, time allocation across different activities, and a brief title for the respondents' primary activity.&#x20;

<details>

<summary><mark style="color:purple;">Survey design, sample selection and response rates</mark> </summary>

The 2-5 minute Typeform online survey was sent out in early November 2023, with a one-week period for respondents to fill out the form and send it back. This form was sent via SMS containing a unique link to the form that helped connect SA Youth registration data to the respondents' survey responses.  Participants received R10 (\~0.6 USD) in airtime vouchers upon successful completion of the online survey.&#x20;

The survey was sent out to a total of 35,000 individuals registered on the SA Youth Platform. of which a targeted 15,000 had been identified from Harambee's EJ data as being involved in microentrepreneurial activities, and 20,000 were new users who had joined the platform within 3 months leading up to November 2023. &#x20;

<mark style="background-color:green;">Total sample:</mark> <mark style="background-color:green;"></mark><mark style="background-color:green;">**35,000**</mark>&#x20;

<mark style="background-color:green;">Total Survey Responses received</mark><mark style="background-color:green;">**: 6,670**</mark>

<mark style="background-color:green;">Response rate:</mark> <mark style="background-color:green;"></mark><mark style="background-color:green;">**19%**</mark>

<mark style="background-color:green;">Total viable responses:</mark> <mark style="background-color:green;"></mark><mark style="background-color:green;">**3,259**</mark>

Of the 6,670 responses received, 1,583 reported having no income-generating activities in the last 30 days, 1,221 provided incomplete/nonsensical answers or were duplicates, and a further 119 were removed due to being older than 35 years.   &#x20;

</details>

#### 5.2.3 Matching descriptions of micro-entrepreneurial activities to ESCO occupations

The next step was to match individual's descriptions of activities from the viable responses with existing ESCO occupations. One option would be to perform this matching exercise manually, similar to the process for the formal economy. However, there were over 3,300 individual descriptions and 3,008 potential ESCO occupations to match these to. Not only would this be significantly labor-intensive, these methods are also highly prone to human-bias and error.&#x20;

<mark style="background-color:red;">\[do we have anything that discusses the risks of carrying over small errors/biases in framework or taxonomy development that can have large consequences in the long run?]</mark>

To streamline the process and mitigate the risk of bias, we opted for a **two-part, sequential matching procedure leveraging artificial intelligence (Ai)**. In the first phase, we designed an algorithm capable of identifying up to three relevant ESCO occupations for each provided description. In the second phase, we reviewed these matches manually to evaluate the accuracy of these matches. Adjustments were made if a match was deemed inadequate or if we believed that an apparent occupation match had been overlooked.

{% hint style="success" %}
**NLP Techniques for Occupation Categorization:** We considered various approaches - including Word2Vec, Doc2Vec, and BERT - but, ultimately, the easiest and most applicable approach for this first exercise was to make a simple prompt to OpenAI’s GPT models. The algorithm used OpenAI's GPT model to analyze all ESCO job descriptions and sort them under the following parameters:&#x20;

* We restricted the potential ESCO matches to the 4.1-digit level of the ESCO taxonomy. Limited scope ensures details while reducing risks of narrow, non-generalizable data on specialized jobs and skills.
* The algorithm could offer multiple categories if there were several potential matches, but it was instructed not to provide more than three matches.&#x20;
* Zero matches were also allowed.&#x20;
  {% endhint %}

In total, 3,236 individuals (99%) were successfully matched to at least one potential ESCO occupation.&#x20;

### Distribution of occupation matched to respondents' answers

{% embed url="<https://datawrapper.dwcdn.net/9Hq4T/1/>" %}

## <mark style="color:blue;">**Activity 3: Assessing relevance of ESCO occupations and skills for informal entrepreneurs via a phone survey**</mark>

Finally, we evaluated the suitability of the ESCO taxonomy for informal microentrepreneurs. We developed a phone survey with three main objectives:

1. **Objective 1:** Verify if the ESCO occupations accurately represent individuals' work and if they would willingly identify themselves with these labels.
   * **Evaluation:** Participants were presented with each matched ESCO label and its description. They were then asked to consider whether these applied to their main way of making money.
2. **Objective 2:** Evaluate the relevance of the ESCO skills associated with each matched occupation
   * **Evaluation:**  W[e randomly selected five essential knowledge and skills and and competences tags associated with each occupation, maintaining proportional allocation between knowledge and skills/competencies tags](#user-content-fn-4)[^4].&#x20;
3. **Objective 3:** Assess the relevance of the identified list of general entrepreneurial skills.
   * **Evaluation:** Five general entrepreneurial skills identified during the previous process were randomly selected for each participant.&#x20;

[**The complete questionnaire can be found here**](https://www.dropbox.com/scl/fi/c4gqb9gq80vv66c4zdbun/IL_Micro-entrepreneurs_Survey2_Questionnaire_v1.0.pdf?rlkey=s9u0wh5qr01ylr28sa5905pg5\&dl=0)

For each skill or knowledge tag, we designed several questions in order to determine its relevance to individuals’ work. These included whether individuals considered the skill as essential or optional, their perceived proficiency in the skill, their perception of how well other entrepreneurs in similar roles perform the skill, the importance of the skill to their own work, its importance to other entrepreneurs engaged in similar activities, and the perceived importance of the skill in a formal sector job akin to their work. We also incorporated a section aimed at identifying potential skill omissions from the ESCO taxonomy for informal micro-entrepreneurs. This section prompted participants to indicate if there were any other skills that were essential for their work that we had not inquired about.

<details>

<summary><mark style="color:purple;">Survey design, sample selection and response rates</mark></summary>

Data collection took place between November 15 and December 8, 2023. The Southern African Labour and Development Research Unit (SALDRU) was contracted to carry out the primary data collection for the phone survey, with Dr. Jacqueline Mosomi as the PI. Data collection was conducted via telephone, assisted by computer-assisted telephonic interview (CATI) software, and the research study received approval from the UCT Commerce Ethics Committee (COM/00513/2023).

&#x20;[**The survey used the sample of 3,236 microentrepreneurs from SA Youth who were previously matched with ESCO occupations**](#user-content-fn-5)[^5]**.** The objective was to achieve 1,500 successfully completed interviews from the specified sample. Respondents were offered a R50 incentive to encourage survey completion, and the survey was designed to take approximately 20 minutes.&#x20;

<mark style="background-color:green;">Total sample: 3,236</mark>

<mark style="background-color:green;">Total Survey Responses received</mark><mark style="background-color:green;">**: 1,515**</mark>

<mark style="background-color:green;">Response rate:</mark> <mark style="background-color:green;"></mark><mark style="background-color:green;">**46.8%**</mark>

[This response rate is comparable to previous phone surveys and exceeds others.](#user-content-fn-6)[^6] The primary reason for non-response (45.8%) was the participant being uncontactable despite multiple attempts. Another 4.1% requested a call back but did not answer, and only 2% refused to participate in the survey.&#x20;

</details>

### Results of the Phone Survey&#x20;

The results of the micro-entrepreneurship survey seem to corroborate our approach, especially by highlighting the absence of systematic bias regarding own skill evaluation between different subgroups in the data. [**The full results of the survey can be found here**](https://www.dropbox.com/scl/fi/1vxk0hdfbmbvbr1b7nrby/IL_Micro-entrepreneurs_Metadata_v1.0.pdf?rlkey=2gfzbjeqhoig5r3okskbdcmvl\&dl=0).&#x20;

<div align="left" data-full-width="true"><figure><img src="/files/cImYPIY4jElxtaFsMptJ" alt=""><figcaption></figcaption></figure></div>

Importantly, young South African on average consider that they are more skilled than the rest of job-seekers. This is positive, as it means that they are confident in their own skills.&#x20;

<figure><img src="/files/wOkEthikG4EIECKJIXkI" alt=""><figcaption></figcaption></figure>

Contrary to what one may expect, the data does not highlight any statistically significant difference in the estimation of respondents' own skills between males and females. If anything, it shows that female respondents tend to deem their own skills as more advanced, compared to their male counterparts.&#x20;

<figure><img src="/files/bZJHU5zbfiA1F9f34fgS" alt=""><figcaption></figcaption></figure>

Finally, there is no statistically significant difference between respondents' own skill evaluation based on whether or not they matriculated (finished high school).&#x20;

[^1]: It is common for young people to a) refer an opportunity that was offered to them to a friend/family in the event that they themselves cannot attend b) bring along a friend/family as a participant. Since turnout on the day of the FGDs was low, the team decided to allow inclusion of these additional members after confirming that they meet all the eligibility criteria for the exercise.&#x20;

[^2]: The final sample included 28 participants from the original sample, along with 5 individuals referred by selected applicants who were not initially chosen from the pool.&#x20;

[^3]: The EU's [Entrepreneurship Competence Framework aka EntreComp](https://joint-research-centre.ec.europa.eu/entrecomp-entrepreneurship-competence-framework_en), describes entrepreneurship as a lifelong competence, identifies what are the elements that make someone entrepreneurial and describes them to establish a common reference for initiatives dealing with entrepreneurial learning.&#x20;

[^4]: ESCO distinguishes essential and optional knowledge, skills and competences in occupational profiles. “Essential” are those knowledge, skills and competences that are usually required when working in an occupation, independent of the work context or the employer, while “optional” may be required or occur when working in an occupation depending on the employer, on the working context or on the country.

    For example, if an occupation had 20 skills and competencies and 5 knowledge tags, we randomly selected 4 skills and competencies and 1 knowledge tag.&#x20;

[^5]: Given that the phone survey relied on the participant being matched to at least one ESCO occupation, it excluded the 23 individuals who were not matched to an occupation.&#x20;

[^6]: For example, just 24% of participants completed the UK Millennium Cohort’s COVID-19 phone survey in February-March 2021 (see <https://cls.ucl.ac.uk/covid-19-survey/content-and-data-wave-3/>).


# Unseen Economy

Tabiya and Harambee's work on the unseen economy is an effort to recognize and utilize human capital built through unpaid work done by young jobseekers, particularly women.

We usually think of "employment" as work that earns money. But many people do important unpaid work at home and in their communities, like caring for family and community members, managing households, organizing community events or volunteering at local organizations. This work helps develop valuable skills. However, these skills are often overlooked in the job market.&#x20;

The problem lends itself well to Tabiya's skills-first approach in the evaluation of human capability. In South Africa, 750,000 young women say they are not economically active due to homemaking activities. However, being a homemaker, does not need to be a sentence. Recognizing the human capital developed in this "unseen" work can unlock pathways into paid employment for these young women and many others like them who are skilled through "alternative" routes. &#x20;

We start by looking at how people spend their time and then [match these activities to skills contained within ESCO](/tabiya-south-africa/unseen-economy/icatus-x-esco-mapping), our skills framework base of choice. This groundwork on taxonomy expansion is relatively straightforward to do, but equally complex to take to market.&#x20;

## Changing Narratives

Our shared understanding of what is and isn't "work", what is and isn't "employable" have been normatively developed over centuries. These perceptions exist both among jobseekers and employers. Skills gained in the unseen economy are not considered as viable or polished as those gained within a seen economy.&#x20;

Changing narratives surrounding skills acquired in the unseen economy is a necessary step both to ensure that job-seekers' are **aware** of skills they have developed, but also that these highlighted skills are **credible** in the eyes of their future employers. Tabiya's works in this space on two fronts:&#x20;

1. [**Compass**](broken://spaces/0O0RbZr6qGsjDH7HKDCo) **for self-discovery**: Many people don't realize the value of skills they've gained from unpaid work. Developed with [support from Google.org](https://blog.google/outreach-initiatives/google-org/google-generative-ai-accelerator-nonprofits/?ref=blog.tabiya.org), [Compass](https://compass.tabiya.org/?ref=blog.tabiya.org) is an AI-enabled conversational tool help job seekers explore and discover all their skills. The tool is particularly useful for this work in that through candid interactions, it facilitates a structured exploration and presentation of people's unseen skills, that many may struggle to think about and articulate in "professional terms".  &#x20;
2. **Research and Advocacy**
   1. **Producing statistical evidence**: Data collected through Harambee's SAYouth platform and various (quasi) experiments are used to build compelling evidence on the quality of skills acquired in the unseen economy, relative to the seen economy.&#x20;
   2. **Advocacy**: By sharing success stories and making scientific evidence understandable to employers, we aim to challenging their biases regarding the skillfulness of job-seekers from the unseen economy.&#x20;

## How Skilled are Unseen Workers ?&#x20;

Including the unseen part of the economy in an usable taxonomy of occupations and skills comes with numerous challenges linked to the difficulty to grasp the skill content of daily tasks and to the diversity of daily habits among young job-seekers. We identified **3 key challenges** associated with skills acquired in the unseen economy:&#x20;

* **Transferability**: Skills acquired in the unseen economy may not be **directly** useable in formal jobs i.e. some certification or skills-testing may be required for transferability.&#x20;

<mark style="color:green;">For instance, many care-takers acquire basic - but also sometimes quite advanced - medical skill if they have been caring for a dependent adult - either and elderly individual or someone made dependent by a physical or mental disability or disease. However, it is very unlikely that those skills are useable in a formal setting, i.e. as a doctor or a nurse, as these two occupations require specific qualifications and trainings under strict regulations.</mark>&#x20;

* **Proficiency**: the level of skill acquired in the unseen economy may not be comparable to the average level of skill of someone performing the same tasks in a formal job.&#x20;

<mark style="color:green;">For instance, a young job-seeker who has acquired cooking skills by talking care of children and feeding them may not have acquired skilled advanced enough - say, cutting techniques - to use them in a high-end restaurant.</mark>&#x20;

* **Likelihood**: performing a certain unseen activity does not imply that the person is using all skills potentially involved or required in this activity, particularly in formal jobs and for complex tasks.&#x20;

<mark style="color:green;">For instance, if a young job seeker declares that he "prepares meals and snacks", it it very likely that they know how to prepare sandwiches and basic dishes. However, it is quite unlikely that they have advanced cooking skills like baking complex deserts - say, an opera cake.</mark>&#x20;

Because of these three concerns, skills acquired in the unseen economy **lack overall credibility, or in  economic terms, signaling value.** One qualifies a skill acquired in the unseen economy as "credible" if employers believe the claim of a young job-seekers declaring that they have acquired that skill in the unseen economy. This issue of credibility is not as forceful in the seen economy as employers are less likely to cast doubts on skills acquired in a formal setup - say, a firm - or in a very identifiable informal occupation - such as informal specialized seller.&#x20;

Addressing the issue of the credibility of skills acquired in the unseen economy - and therefore issues of transferability, proficiency, and likelihood - is essential to Tabiya's goal of changing narratives around unseen activities. On the one hand, our task is to highlight the skills acquired in the unseen economy. On the other hand, acting as if skills acquired in the unseen economy were as credible as skills acquired in formal jobs would do disservice to young job-seekers from the unseen economy. Indeed, employers ultimately take hiring decisions based on their own beliefs and perception of the credibility of skills.&#x20;


# ICATUS X ESCO Mapping

The ICATUS Taxonomy provides a complete picture of how people use their time. Based on the assumption that all activities build human capital, we attempt to connect ICATUS and ESCO.

## About ICATUS&#x20;

The International Classification of Activities for Time-Use Statistics (ICA-TUS) classifies ["all the activities on which a person may spend time during the 24 hours that make up a day"](#user-content-fn-1)[^1].&#x20;

ICATUS activities are meant to partition the universe of conceivable daily activities in an exclusive manner: each activity can only fall into only one ICATUS category. Therefore, the taxonomy provides a solid starting point and strong conceptual structure into our work on the [unseen economy](/tabiya-south-africa/unseen-economy). Additionally, as the universally-recognized and accepted standard framework for time-use statistics, it allows for nation-wide and cross-country comparisons of how people use their time.&#x20;

However, while ICATUS is designed to be consistent in definitions with economic frameworks like the SNA 2008 and ISIC, it serves a broader purpose than defining economic or productive activities. Thus, it is not directly mappable to the ISCO, and by extension, ESCO skills - the core framework we use.&#x20;

Hence, the first objective of Tabiya and Harambee's work on ICATUS was to match its activities to ESCO skills. In the long run, in concept, this could allow SAYouth users, from the unseen economy or with experience in the unseen economy, to select their main time-uses and be presented with candidate skills to highlight on their CVs and inform a job-matching algorithm.&#x20;

## Assigning ESCO Skills to ICATUS Activities

The teams had two levels at which these matches could be made:&#x20;

1. **Time-use activity to skill:** Matching ICATUS activities directly to ESCO skills using natural language processing (NLP). In this case, skills matched may be spread out across various occupations and not cohesive. Matching outcomes and precision would be systematic but highly reliant on the richness of skills and activity descriptions.    &#x20;
2. **Time-use activity to occupation:** Matching ICATUS activities to ESCO occupations. In this case, we would be tentatively attributing all skills under the matched occupation as a potential candidate skill. Given the shorter list of occupations to match to, this could be done manually by team members. &#x20;

Given the scarcity of the readily available data on platform users' description of their daily tasks, the second option was selected. ICATUS activities were thus each matched to up to 4 ESCO occupations manually by a team of 3 researchers. These matches then allowed the team to derive a candidate list of skills for each ICATUS activity by the ESCO taxonomy.

Assigning ESCO occupations to ICATUS activities was based on similarity between their definitions. For instance, "preparing meals and snacks" in ICATUS involves - according to its definition - tasks that are those of "cooks" and "kitchen assistants" in ESCO. Therefore, the complete list of ESCO skills associated with "cooks" and "kitchen assistants" was assigned to "preparing meals and snacks" in ICATUS.&#x20;

<figure><img src="/files/WLub9KGLM1UxmhbSGvGf" alt=""><figcaption><p>Demonstration of a first-stage ICATUS X ESCO mapping for a commonly reported time-use activity: "Preparing meals and snacks" </p></figcaption></figure>

## Trimming the Candidate List of Skills

The skills list obtained from the first-stage mapping was hardly useable. Each ICATUS activity ended up being associated with a large number of skills. In the example pictured above, "Preparing meals and snacks" generated a first-stage list of 59 skills. These skills fall along a broad spectrum of signaling power and relevance for a jobseeker from the unseen economy: some are too elementary to be worth mentioning such as "washing and cleaning vegetables" while others being too formal-sector specific for most home-cooks to have relevant knowledge of such as "designing indicators for food waste reduction".&#x20;

In order to provide a manageable list of skills to job seekers on the [Harambee Platform 0.1](/tabiya-south-africa/harambees-platform-version-0.1), the second-and third stage of this exercise attempts to trim this tentative list of skills. First, the Tabiya team removes skills deemed obviously not consistent with the definitions and description of the respective ICATUS activities. This was done to provide a relatively leaner and relevant list to experts at the [Harambee panel of South African labor market experts and intermediators](/tabiya-south-africa/unseen-economy/harambee-panel). &#x20;

For this second stage skills selection, Tabiya followed 4 rules: &#x20;

#### <mark style="background-color:yellow;">#1 Isolate Knowledges</mark>

**As a first step, we isolated** [**knowledges** ](#user-content-fn-2)[^2]**and mainly focus on skills,**  to the extent that certain knowledges listed in ESCO tend to be redundant with skills i.e. their presence in the list does not bring new information content. <mark style="color:orange;">For example, if a young job seeker "uses cooking techniques", this implies that they to know "cooking techniques".</mark>&#x20;

However, certain knowledges were kept in for panelists' discretion when they brought important new information. For instance, "common children's diseases" may be deemed an essential knowledge for child care, whether or not one has firsthand experience of having dealt with them.&#x20;

<mark style="background-color:yellow;">**#2 Delete irrelevant skills:**</mark>

**We exclude skills that had been unduly associated to ICATUS activities** on account of the two frameworks not being completely aligned in definition, concept and function. In ICATUS, each activity is clearly defined, including a detailed list of tasks that are included or excluded, along with examples. ESCO follows a similar structure, but ICATUS activities are usually broader and conceptually different from ESCO occupations.&#x20;

<mark style="color:orange;">"Budgeting, planning, and organizing duties and activities in the household", an ICATUS activity, is not an explicit occupation and thus absent in ESCO. To align ESCO skills with this ICATUS activity, it has been matched to the ESCO occupations "Office Clerk," "Accountant," and "Bed and Breakfast Operator," which cover tasks related to household management. However, each occupation extends beyond the ICATUS activity. For example, "Bed and Breakfast Operator" includes the skill "serve beverages," a task that is completely irrelevant to the ICATUS activity in question.</mark>&#x20;

<mark style="background-color:yellow;">**#3 Delete skills that are deemed too formal**</mark>

Some skills inside ESCO occupations directly refer to situations or tasks only possible to do in the seen economy, whether it is formal or informal. <mark style="color:orange;">For example, "maintain customer service" cannot be appropriately associated with "serving meals and snacks", as a home-cook would not have "customers" and would only be serving meals and snacks to one's own family members or at best house guests.</mark> The underlying assumption here is that this process requires using only literal meanings of skills; in this case,  unpaid service to family and guests is fundamentally different from having to serve paying customers.&#x20;

<mark style="background-color:yellow;">**#4 De-duplication**</mark>&#x20;

Duplicated skills, both verbatim and in content, were removed from the aggregate list of skills originating from stitching all potential ESCO occupation matches. <mark style="color:orange;">For example, "Outdoor cleaning" was associated with both "prune plants" and "prune hedges and trees", it is debatable how much marginal information retaining both provides.</mark> &#x20;

However, content de-duplication is not straightforward within ESCO. <mark style="color:orange;">In the above example, "prune plants" does not contain the subskill  "prune hedges and trees", even though hedges and trees are obviously plants. Instead, it contains the subskill "perform hand pruning".</mark> Basically, the organization of ESCO itself is not straightforward or intuitive, complicating content-based de-duplication.

##

[^1]: <https://unstats.un.org/unsd/demographic-social/time-use/icatus-2016/>

[^2]: The ESCO classification distinguishes between "skills" - that describe an action - and "knowledges" - that describe know-how. For example, the occupation "cook" is associated with the skill "use cooking techniques" and the knowledge "cooking technique".&#x20;


# Harambee Panel

As one of our main tools to build a localized and inclusive ESCO taxonomy for South Africa, Tabiya organized a panel discussion to channel Harambee's knowledge of South African young jobseekers.

{% hint style="warning" %}
The panel ended in May 2024 and its results are still being processed. This page is thus a work in progress.&#x20;
{% endhint %}

When it comes to assessing the transferability and credibility of skills acquired by job-seekers in the unseen economy, employers are likely to display strong biases. Namely, when a young job-seekers declares that he possesses skill A, an employers is likely to consider that skill A acquired in the unseen economy is not equivalent to skill A aquired in the seen economy, both in terms of the level of skillfulness (credibility) and nature of the skill (transferability).&#x20;

Employers hold strong biases and priors when assessing the transferability and credibility of skills acquired by job-seekers in the unseen economy. That is, when a young jobseeker claims that he possesses skill A, an employer will likely note that skill A acquired in the unseen economy is not equivalent to skill A acquired in the seen economy, both in terms of substance (**transferability**) and intensity (**credibility**). This is a key challenge when building a taxonomy inclusive of unseen experience.&#x20;

To build an inclusive taxonomy, our challenge is, therefore, to try to get an unbiased overview of the skills that could be associated with unseen activities defined in the ICATUS framework. There exist no data systematically estimating the skillfulness of workers and people involved in the unseen economy, although some work has been conducted to assess the acquisition of basic skills (literacy, numeracy) among Harambee users (cf. Kate Orkin). Our approach therefore relies on an attempt to oppose different biases in order to get a sense of the credibility of skills associated with unseen activities. Our guess is that, by making people who have different views on the unseen economy discuss, we may obtain a clearer picture of the value agents give to skills acquired in the unseen economy.&#x20;

## Panel Participants

Representation across 4 key interests/stakes in this space are ideal for a panel with these objectives:&#x20;

1. **Employers.** Understanding the demand-side perceptions on credibility of skills acquired in the unseen economy is essential to building an intermediation tool that successfully promotes the skills of jobseekers coming from outside the labor market. &#x20;
2. **Intermediation experts** who work with both supply (jobseekers) and demand (employers) sides reflect both the will to improve labor market inclusivity and realistic understanding of what works for job market success.&#x20;
3. **Jobseekers from the unseen economy**  can provide the clearest idea of what tasks they are capable of and to what extent. Additionally, they can also provide valuable insights on tendencies and upward biases in overreporting and misrepresenting skillfulness to get jobs. &#x20;
4. **Auxiliary beneficiaries of the unseen economy** are all individuals that benefit from unseen work including children, dependent and non-dependent adults, families operating household firms etc. These participants can offer valuable third-party perspective on skills acquired in the unseen economy by jobseekers.

Mobilizing such a large and dispersed panel of individuals is complex and difficult to implement. Specifically, it was deemed close to impossible to mobilize employers to participate in such an exercise within a reasonable timeframe. Looking back after completing the exercise, it would have indeed been extremely difficult given the length of the exercise. As a result of the different constraints, the panel was held with 5 to 7 members of Harambee, the number of panels percipients varying between the phases of the exercise and in-between online surveys. While the limited panel participants made it difficult to aggregate answers, it made the exercise more manageable, especially when it came to ensuring that responses from panelists within the given timeframe and ensuring attendance at the online panel discussion. All that being said, the experience gathered during the Harambee panel may help us to bring to the table more diverse actors in the future.&#x20;

## Methodology of the Panel&#x20;

~~The OMS research team’s initial proposal for the organization of the panel was not possible to implement due to the lack of time of panelists and the overall length of the exercise. Adapting the exercise to the constraints gradually discovered by the team allowed it to come up with a more realistic methodology that is replicable, albeit improvable. This document presents the new methodology used after the initial conception of suggested skill lists by the OMS team.~~

Participants were first briefed on key concepts to be used throughout the panel, followed by an online survey that aimed to elicit participants' perceptions/assessment of signaling values of skills under specific activities. The last phase was an in-person panel discussion focused on transferability and proficiency.&#x20;

### 1. Concept briefing

In panels like these, that attempt to bridge theoretical interpretations of labor market behavior with practical understanding of the space, i.e. a collaboration of academics and practitioners, it is important that all parties precisely agree on the terminology and concepts being referenced. Panel participants were first provided a briefing on concepts used to describe the signaling value of skills and experiences, to be read before starting the online surveys.&#x20;

<details>

<summary><mark style="color:blue;">Concept brief provided to participants</mark></summary>

We now have a list of ICATUS Activities mapped to candidate ESCO Skills. We need to solve for the following issues before taking this work to platform (version 0.1):&#x20;

Meaningfully reduce the number of skills we assume a participant of the unseen economy  may have based on the signaling value of these candidate skills to potential employers. We want to do this in a way that mitigates biases and changes problematic human capital valuation narratives.&#x20;

**By signaling value, we mean:** Understand which skills are a meaningful signal to employers. We define signaling as a function of the employer perceived transferability of a given skill from the unseen economy to a formal sector occupation in the seen economy; likelihood that a given skill is actually used in a given ICATUS activity; and the proficiency level in a skill that employers think participants may have acquired relative to seen economy workers. The weight of each of these signaling function variables requires empirical assessment for the South African context. Moreover, this signaling function may include other factors not mentioned here.&#x20;

**By transferability, we mean:** the relative ease of applying a skill gained in the unseen economy to an occupation in the formal labour market. A skill is relatively more transferable if the skill can be applied the same way in the unseen economy and in the formal labour market. An unseen economy skill is relatively less transferable if the skill requires intervention such as upskilling to be applied to the formal labour market.

**By likelihood, we mean:** the perceived chance that an unseen economy participant uses and possesses a traditionally seen economy ESCO skill. That is, assuming the same activity were to be performed in the unseen economy (an environment with less formality and regulation on average) and the seen economy, which skills would likely not be used at all in the unseen economy context.

**By proficiency, we mean:** how well we assume an unseen economy participant can use a skill derived from the unseen economy relative to their counterpart in the seen economy, assuming that the skill is perfectly transferable and it is perfectly likely that an unseen economy participant possesses this skill.&#x20;

Importantly, this signaling-based reduction of candidate skills must mitigate the persistence of biases that adversely impact the unseen economy job seekers' labour market outcomes. By biases, we mean that one can expect potential employers and unseen economy participants to systematically under and/or over estimate the transferability, likelihood and proficiency levels of the skills in question in this work. Therefore, an intermediary between job seekers and employers is best placed to perform this exercise.

</details>

### 2. Online Survey

The online survey had 6 modules covering all ICATUS activities under the three categories for unseen economy: **Group 3: Unpaid domestic services for household and family members, Group 4: Unpaid caregiving services for household and family members,** and **Group 5: Unpaid volunteer, trainee and other unpaid work.** In each module, [\~8 ICATUS activities were presented to participants](#user-content-fn-1)[^1], and for each the activities, all suggested skills were shown. For ease of use, the research team pre-organized the skills into categories similar skills, such as “communication skills” or “management skills”. For each skill, respondents were prompted to assign a signaling value score ranging from 0 “no signaling value” to 3 “high signaling value”. By the end of the exercise, a total of 7 responses were collected for the first two modules, 6 each for modules 3, 4, 5 and 6.&#x20;

<figure><img src="/files/AysqPHJswn9Rp7yX8MIW" alt=""><figcaption><p>Example of overall prompt for an ICATUS activity</p></figcaption></figure>

<figure><img src="/files/SpPANKcEeChrfFEX25hW" alt=""><figcaption><p>Example of prompt for a skill category</p></figcaption></figure>

### 3. Virtual Panel Discussion

The Harambee in-person panel was the final step in associating a final list of ESCO skills to ICATUS activities, using signaling value as a discriminating characteristic. The two- and half-hour session included all online survey participants except one. For this discussion, the research team first identified skills that demonstrated high variance in signaling value based on the survey results. On the panel, participants were asked to discuss the signaling value of these high-variance skills, first focusing on on likelihood, transferability, and proficiency, before opening discussions into other considerations.&#x20;

## Results of the Panel

The Harambee panel helped achieve clarity on several aspects of this unseen skills-association exercise: first, it allowed the research team to adapt the list of unseen activities to the South African context; second, it allowed the research team to better understand concepts used by labor market intermediation professionals in thinking about the signaling value of skills. The exercise was also a useful proof of concept in combining qualitative inputs to taxonomy building and suggests that reproducing this methodology in other national contexts will be essential.

### Online Survey Results

For each activity in ICATUS, skills mapped to it were scored by up to 7 Harambee participants on the online survey based on their perceived signaling value. There were two options in aggregating these scores: we could average the results or assign final value based on majority results. Given the small sample of responses, both of these approaches had high risks of generating "spurious precision". Therefore, the team opted for a segmented approach based on variance in perceived signaling value:&#x20;

1\) scores for skills-occupation dyads with no consensus (i.e. half 0 “no signaling value” and 3 “high signaling value”) were set aside for further deliberation through the in-person panel.&#x20;

2\) low variance skills-occupation dyads were assigned a final score that summed all responses.&#x20;

<details>

<summary>Cleaning and aggregating online survey results </summary>

Between March 15, 2024, and May 17, 2024, we conducted six Google Forms surveys with Harambee panel members to determine a signaling value score. Answers were calibrated with "no signaling value" as zero, "low" as one, "medium" as two, and "high" as three. This calibration facilitated visualization of each survey's distribution and helped identify skills with the highest variance in responses, indicating divergent perceptions among panelists.

For coherence, Surveys 1 and 2, associated with ICATUS Major Division 3, were combined, excluding at most two Group Level Activities from Divisions 3 to 5 that had minimal impact on distribution shape. Similarly, Surveys 3 and 4 were combined under Major Division 4, and Surveys 5 and 6 under Major Division 5, streamlining the analysis. Surveys 1 and 2 had seven responses, Surveys 3 and 4 had six, and Surveys 5 and

</details>

### Learnings about subjective signaling value functions

Despite the initial instruction for the panel to use likelihood, transferability, and proficiency as the signaling value pillars to score skills by, there was frequent debate about other components that the panel included in the scoring. The first and most prevalent other component was South African context specific considerations. The panel often relied on their understanding of the South African labor market to assign signaling value scores. In fact, they relied on this as much (in frequency terms) as the other likelihood, transferability, and proficiency indicators. The other concept that frequently came up during the panel was the panel’s expertise on the expectation of the response of the average South African entry-level employer to a given skill. Given the Harambee panel’s labour market expertise, they were particularly well placed to infer how a representative employer would react to a skill. Thus, the panel used this inference in deciding which high-variance skills to assign no, low, medium, and high relative signaling value strengths to.

The virtual workshop provided the first opportunity for panelists to engage live with one another over the survey and skills selection process. Hence, two ad hoc decisions we made by the panel. First, the panel elected to remove the term “social service users” from ESCO skills. Due to the term’s ambiguity and incompatibility with the South African context, the panel elected to have it replaced with a more intuitive term such as “people” or “dependents”. The final decision implemented by the research team has been to replace ESCO skills comprised of the term “social service users” with similar ESCO skills that do not include the term. For example, the ESCO skill “Assist social service users with physical disabilities” will be replaced with the ESCO skill “assist individuals with disabilities in community activities”. This rule has been adopted for uniformity in all skills in the localized taxonomy related to social service users.

In addition, the ICATUS Group Level activity entitled “Tending furnace, boiler, fireplace for heating and water supply” was removed from the localized taxonomy because of the inapplicability of the activity and its associated skills to the South African context. The closest version that could replace “Tending furnace, boiler, fireplace for heating and water supply” that is found in the South African Time Use Survey at the Group Level is “Chopping wood, lighting fire and heating water not for immediate cooking purposes”. Despite the existence of this potential replacement, the panel elected that this replacement was not worth including in the finalised taxonomy due to its likely inapplicability and scarcity in the South African labour market.

During the virtual workshop, the panel expressed that it found transferability, likelihood, and proficiency to be good signaling value indicators. They found that these indicators were especially enhanced with their knowledge of the South African context. The panel members cited lengthiness and repetitiveness as the main criticisms of the process. Adding page numbers to each survey was suggested as a way to manage survey length expectations and, in turn, better allocate time to completing the survey.

#### &#x20;Types of arguments

a.        Likelihood

The argument of the likelihood that a jobseeker has a skill was used very frequently during the discussions. The arguments relating to the likelihood of a jobseeker having a skill were based varied among examples. In most, they were based on the knowledge respondents have about the South African context. For instance, “teaching young horses” was deemed unlikely in link with the “daily pet care” activity, with respondents acknowledging that they would expect the likelihood to be higher in a European context. Another example is that driving skills were deemed to have a low likelihood in link with activities consisting in accompanying dependent adults or children, as participants stressed that moving around entailed less car rides than it would in Europe. In some other cases, the likelihood that a jobseeker has a skill based on an activity is deemed to be low because the skill tag is conceptually remote from the activity title. For instance, although a young jobseeker volunteering in a church could possibly participate in accounting activities, accounting skills seem to be conceptually to far remote from the idea an employer may have of volunteering in a church. Finally, respondents stressed that to answer they had to assume the exact activity young jobseekers were doing. For instance, volunteering in another household’s firm is not precise enough for an employer to infer what skills the jobseeker may have. This calls for a further improvement of the unseen taxonomy.

b.        Transferability

Transferability was another of the main concepts presented to panel participants. Therefore, the concept was also widely used during the conversations. Sometimes, the concept was used in the right way. For instance, preserving food has been deemed to be weakly transferable, as professional food preservation relies on scientific principles and rules that one does not typically use in the unseen economy. However, the concept of transferability was also mixed up with other concepts multiple times. For instance, it was mixed up with likelihood, when participants stated that a given experience “did not transfer” to a skill. It was also used concomitantly with proficiency. When is skill was deemed to be basic, for instance tending grass, then it as deemed transferable. Overall, panel users seemed to deem skills to be on average more transferable than likely, but the two concepts appeared as equally important.

c.        Proficiency

The concept of proficiency was almost not used by panel participants. More precisely, it was only invoked through the transferability of skills, when participants guessed that basic skills were more transferable and thus had more signalling value. Otherwise, discussions around how good a jobseeker might be able to get at given skills through given activities was not discussed.

d.        Labor market demand

Discussions around labour demand were rather unwanted by the team prior to the panel, as the research team hoped that “what employers want” would not influence the making of skills lists for ICATUS activities. However, some interesting conversations happened where panellists stressed that some unseen experiences could have a very signalling value for very specific employers. They thus deemed it important to keep certain skills visible on our skills lists, to allow jobseekers to signal them. On the contrary, some skills were left aside as participants deemed the demand for it to be low.

e.        South African context

As described before, the South African context was mobilised multiple times along with the concept of proficiency, usually to argue that a skill or occupation was not relevant to the South African context. The later was unexpected, as the panel decided to delete “Tending furnace, boiler, fireplace for heating and water supply” and “Accompanying non-dependent adults” from the list of unseen activities altogether. Considerations around the South-African context fed into discussions about other concepts such as likelyhood and transferability, for instance through qualifications. Indeed, some skills were deemed not to be transferable on the ground that they require qualifications in South Africa.Such an account for the South African context is perfectly aligned with what the research team hoped the panel would discuss.

f.           Stretch between the activity title and the skill

Participants stressed multiple times that although skills relating to an activity might be likely and transferable, they are sometimes too far remote from each other. Therefore, an employer would be unlikely to think about skill A if a young jobseeker came with the unseen experience X. Conversely, some individual skills were deemed very valuable in themselves, but the panel thought that it is unlikely that a user picking the activity they are associated with would do it to pick them. Namely, some skills are unlikely to be picked, not because they are not valuable or transferable, but because picking them would require picking activities that users are unlikely to choose. This finding was of great relevance for the work of Tabiya, as it suggests that a taxonomy like ESCO may not be flexible enough for jobseekers to surface their human efficiently.

g.         Naming of skills

Finally, multiple comments from participants referred to the naming of skills, that was deemed to make skills un-pickable for young jobseekers. In particular, many skills refer to “social services users”, for instance “assist social service users with physical disabilities” or “social service users to live at home”. Although the human capital signalled by these skill-tags was deemed to be relevant by panellists, they noted that the added precision of “social service users” made the skill less relevant, as it refers to a much more formal environment than the ones young jobseekers from the unseen economy are working in. In practice, skills like “assist social service users with physical disabilities” were deemed to have a low signalling value, although it was recognized that a skill like “assist individuals with physical disabilities” would have received a higher score. Yet, these skills do not have neutral equivalents that do not refer to “social service users”. Therefore, it was suggested to create new skills.

#### Takeaways about the signalling value function

The in-person panel proved useful to understand how Harambee workers think about the signalling value of skills in relation with occupations.  Although we imposed three concepts (likelihood, proficiency, transferability), panellists appeared to find them useful, although proficiency was discussed less than the two others. Interesting remarks that feed into our concept of signalling value include the importance of qualifications (a skill may have a proficiency, transferability, and likelihood score but send a weak signal because it typically entails formal training). Second, it was interesting to see that skill tags influence the signalling value scores when they become too specific to a given context (social service users).&#x20;

## &#x20;Learnings for future country-specific taxonomies

The in-person panel served as a proof of concept for further use of this qualitative methodology to adapt ESCO to local taxonomies. The exercise was deemed successful overall, albeit time-consuming. Unexpectedly, the exercise also brought about discussions that seem to call for solutions like Compass to overcome the rigidity of taxonomies.&#x20;

### Iterations of the panel exercise

Discussions surrounding the South African context, rather they touch upon what unseen activities young jobseekers might do or upon qualifications and demand side needs brought important insights used in building the “unseen” part of our localized taxonomy. Therefore, it gave encouraging results when it comes to using the same methodology in our next countries (France, Ethiopia…).

When it comes to organizing the panel, the research team ended up changing its initial plans multiple times. Online surveys proved to be very time consuming for respondents, who explicitly suggested that the research team shortens lists more than it did already to facilitate the work of panellists. The in-person panel on the other hand seems to have been liked by panellists and was fruitful. The suggestions for further panel iterations are the following:

1 – Apply the skills selections rules more stringently before presenting them to the panel, to reduce the time needed to answer the online surveys.

2 – Organise an in-person panel session to answer the first survey, along with the training session that was provided in this iteration. This was suggested by panellists so that the variance in answers is lower.

3 – Panellists should initially be asked which of the ICATUS activities they deem relevant for their local context, to avoid the situation of “tending furnaces”.

4 – Online surveys should include an indication of the overall length of the survey.

### Compass

It is to be noted that this panel exercise suggested the need for a solution like Compass. Panellists stressed multiple time the lack of connection between skills and the ICATUS activities that they were linked with. While recognizing that some skills could have a high signalling value and be very relevant in the South African labour market, they considered that young jobseekers were unlikely to be able to highlight them because they were unlikely to pick certain ICATUS activities. This highlights multiple sources of rigidity in localized ESC0 taxonomies. First, our taxonomy is based on ICATUS, which covers all productive non-paid time uses but groups them in buckets that may not speak to platform users. Second, UX constraints call for the lists of skills associated to each ICATUS activity to be limited, which necessarily leaves aside certain skills. Therefore, the unseen taxonomy does not include all the skills that a young jobseeker might want to highlight, which is a common issue with ESC0 and results from the trade-off between a taxonomy that is inclusive (includes many skills for all occupations) and informative (limits the number of skills associated with each occupation). One solution to this is to have a flexible algorithm, allowing jobseekers to find skill through their experiences but also *ad hoc*, and thus highlight the full range of their human capital.

## Online Survey Results

### Distributions

Each survey’s distribution is shown below. Each ICATUS Major Group has relatively different distributions. This is likely driven by the fact that the survey respondents’ sample size changed within each major division in an already small panel population pool. Note that the signal values are the aggregation of individual panel respondents’ answers. Therefore, Figure 1’s signal value axis is larger than Figures 2 and 3, which each had fewer respondents.

<figure><img src="/files/wsFo8zNvR2oeBp64IlDj" alt=""><figcaption><p><em>Figure 1 – ICATUS Major Division 3 Distribution</em></p></figcaption></figure>

<figure><img src="/files/S8L2u1vKc8QuTUS0eJSD" alt=""><figcaption><p><em>Figure 2 – ICATUS Major Division 4 Distribution</em></p></figcaption></figure>

<figure><img src="/files/1hzxt5wGxi8f6LP78g3A" alt=""><figcaption><p><em>Figure 3 – ICATUS Major Division 5 Distribution</em></p></figcaption></figure>

### Variances

The variance of answers associated with each skill was particularly important because of their applicability to the live panel workshop discussion. The rationale used was that skills with the highest signaling score variances would be the most worthwhile to attain agreement on amongst all the survey respondents, given that this exercise could only be performed on a subset of all the skills chosen during the survey.

We isolated 20 of the skills with the highest variances per ICATUS Major Division. Ultimately, though, the skill with the 20th highest variance within each Major Division shared this variance with at least four other skills. Thus, the list of 20 high-variance skills expanded in each Division. A total of 26 skills ultimately comprised the high-variance skills for ICATUS Major Division 3, 24 for ICATUS Major Division 4, and 47 for ICATUS Major Division 5.

Following the virtual workshop, the high variance skills that were not discussed were assigned inferred signal scores based on the learnings from the panel. These scores were sent to three Harambee panel members for verification before their final incorporation into the taxonomy.

#### Virtual Group Discussion Results

A virtual two-and-a-half-hour workshop to discuss skills with the highest occurring variances per survey was hosted on the 30th of May 2024. Six of the seven survey respondents joined the session (though some respondents had to move into and out of the workshop). From this session, 12 skills from Surveys 1 and 2 (ICATUS Major Division 3) were discussed, nine skills from Surveys 3 and 4 (ICATUS Major Division 4) were discussed, and nine skills from Surveys 5 and 6 (ICATUS Major Division 5) were also discussed (this does not include skills and activities which need rewording or will be deleted from the taxonomy).&#x20;

Of the 12 ICATUS Major Division 3 skills discussed, 83.3% were deemed to have no signaling value (a signaling score of zero), none of the skills were deemed to have low signaling value (a signaling score of one), 16.7% were deemed to have medium signaling value (a signaling score of two), and none of the skills discussed obtained a high signaling value. For ICATUS Major Division 4, 55.6% of the total ten skills were assigned a no signaling value score (a signaling score of zero), none were assigned a low signaling value (a signaling score of one), 22.2% were assigned a medium signaling value score (a signaling score of two), and 22.2% of the skills discussed obtained a high signalling value score. Finally, 66.6% of the nine ICATUS Major Division 5 obtained a no signaling value score (a signaling score of zero), and the remaining 33.3% was equally shared between low, medium, and high signaling value scores amongst the discussed skills. Figure 4 depicts the workshop’s signal value discussion over all ICATUS groups. It is clear from the figure that skills that were heavily contested in the survey were ultimately deemed to have no signalling value once panel members had a chance to pose arguments for and against these skills.

<figure><img src="/files/uJcwv4U9RpF5zZZSOSIZ" alt=""><figcaption><p><em>Figure 4 – High Variance Skills Signal Value</em></p></figcaption></figure>

### Final Skills Selection

Once we received the remaining high variance skills’ scores, we adopted a “traffic light” decision rule. That is, 33.3% of the skills with the lowest signaling value scores within an ICATUS Group (highest level of disaggregation) form part of the red category. These skills are removed from consideration on the SA Youth platform. 33.3% of the skills with the highest signaling value scores form part of the green category. These skills are surfaced at the top of the list of skills options that SA Youth users can select. The mid-range signaling value scores comprise the remaining 33.3% and are in the orange category. These orange category skills are surfaced beneath the full list of green category lists on the SA Youth platform. Below are three diagrams that show the absolute number of skills that will remain due to the enforcement of the traffic light decision rule for each ICATUS Major Division.

<figure><img src="/files/3pXMBs5ossLhLmzUSkrY" alt=""><figcaption><p><em>ICATUS Major Division 3 Skills Split</em></p></figcaption></figure>

<figure><img src="/files/zmlwC1X6IQzY4GLeDuRc" alt=""><figcaption><p><em>ICATUS Major Division 4 Skills Split</em></p></figcaption></figure>

<figure><img src="/files/Im9Lj4NLogwr2Ni6z8Dq" alt=""><figcaption><p><em>Figure 7 – ICATUS Major Division 5 Skills Split</em></p></figcaption></figure>

[^1]: Within each ICATUS category, we excluded activities that are considered “other”, for instance “other unpaid domestic service for household and family members” as the skills and activities are identical to the rest within each ICATUS category.


# Harambee's Platform Version 0.1

Tabiya's work on the South African taxonomy will be directly channeled by Harambee through the SAyouth.mobi platform. Concomitantly, it is used as a basis for Compass, our AI chatbot.

The optimal way to use our inclusive taxonomy to help young jobseekers is not obvious. For instance, it could be used as the basis of a static survey in which users pick occupations and skills manually, or in a more flexible way by using AI. Although certainly complementary, we hope to test these two approaches separately to identify the advantages and disadvantages of each approach.&#x20;

{% hint style="warning" %}
Harambee and Tabiya work closely to come up with the version 0.1 of the Harambee platform. As the final product is highly dependent ongoing research work, this section will be filled later.&#x20;
{% endhint %}


# Ordering of Skills: Review of Evidence

There is ample evidence that the way options are presented to users influence their picks. Understand users' biases is critical to improve the functioning of the Harambee platform.

Experience has shown that young jobseekers struggle to identify skills that are the most relevant to their professional occupations and hobbies, or the skills at which they are the most proficient. Yet, helping them navigate ill-functioning labour markets entails surfacing the full length of their human capital – namely, the full set of skills they possess -, while helping them to present it in a strategic manner to improve their job seeking perspectives. To help jobseekers do so, Harambee hopes to rely on static surveys prompting young jobseekers to select, for each of their occupation / activity, the five skills that they are the most proficient at. To do so, they will be presented with lists of ESCO skills associated with their activities, or with skills that have been selected through the work of Harambee and the Oxford Martin School (for micro-entrepreneurship and unseen activities). This raises the question of the optimal way to present skills to job seekers, to ensure that 1 – we cover their whole human capital, 2 – they can effectively pick relevant skills to populate their inclusive CVs, 3 – this process is efficient. Yet, evidence has shown that the way individuals are prompted to pick answers out of list of options impact their final decisions. In this note, I summarize some research pieces to inform how the lists of skills submitted to jobseekers by Harambee and OMS may influence the decisions.

## Choice Overload and Underload

These notes are based on « Cognitive and Affective Consequences of Information and Choice Overload” by Reutskaja et al.

Harambee has expressed concerns that the number of options presented to job seekers may overwhelm them, making their experience on the platform anxiety inducing and the results inefficient. These concerns are aligned with insights from behavioural science, focusing on the concept of “choice overload”. Intuitively, individuals presented with too many options to pick from may find it harder to pick ad experience anxiety, unsatisfaction, loss of motivation to pick, stress, or even choice paralysis. It is to be noted nevertheless that the reverse phenomenon, choice underload, may lead to frustration as well, as the desired options may not be presented to individuals. Behavioural theory suggests that the level of satisfaction felt by individuals prompter to pick a choice out of several options follows an inverted U-shape. Namely, the optimal number of options to give individuals is not the maximum.

<figure><img src="/files/cfYGK38r7uGTZ4h84qnT" alt=""><figcaption></figcaption></figure>

&#x20;          The intuition behind this is the following. When presented with an increasing number of options, the needs of people are more likely to be satisfied. In our work, more options means that Harambee’s young job seekers are more likely to find the skills that they are the most of proficient at among the ones we present them with. This abundance of choices can also be beneficial from a psychological point of view. In our case, it may comfort users into the idea that they have large sets of skills, and thus allow them to have a more positive approach in their job search. On the other hand, large sets of options have psychological and time costs. Studies have shown that our minds do not account for the full sets of available information when taking decisions, but only subsets (Wedel and Pieter, 2007). Following Reutskaja et al’s terminology, this “inattentional blindness” means that jobseekers presented with long lists of skills risk ignoring skills that they would have picked in smaller sets of options. In the case of Harambee users, a large set of options may also induce stress, especially linked to the sentiment of forgoing important skills. Worse, it may lead to a loss of motivation to answer the prompt correctly, and to users selecting skills by default. This evidence suggest that Harambee is right to worry about the number of skills presented to young jobseekers being too large.

## Ordering of Skills

Once the lists of skills presented to young jobseekers have been narrowed down, the order in which the skills are presented is likely to impact their picks. As Harambee and the teams in Oxford are considering ordering the skills based on concepts such as transferability (essentially, the frequency with which a skill is associated with an occupation in ESCO), understanding how that order may influence users’ pick is critical. Intuitively, one could think that users tend to pick the first skill presented to them, especially when they must take an action (click “see more”, for instance) to see other options. Studies researching the effect of the order of questions in online surveys suggest that order does matter. A study conducted in the US in 2008 found that low-education respondents who answer surveys quickly – which is likely to be the case of Harambee users – a most prone to primacy effects (bias toward selecting earlier response choices). A study conducted in Xiao et al in 2024 also observes such bias, while finding that it is inexistent in “objective questions”, for instance those pertaining to demographics.

The implications for the Harambee platform are not straightforward. On the one hand, Harambee may seek to nudge job seekers into picking skills that Harambee’s counsellors consider as better picks. This would imply putting these skills on top of the list of options. On the other hand, Harambee may choose not to try and influence users’ picks. However, a random order would effectively lead to answers biased toward the first options presented. Countering this bias is not evident, which may be another argument to choose to hierarchize skills.&#x20;

{% hint style="info" %}
*Nudging and Agency:* the paradox behind libertarian paternalism is at the heart of the criticisms addressed to R. Thaler and C. Sunstein’s 2008’s book *Nudge*. Their argument is that modifying the architecture of choices – namely the way the information is presented to individuals bound to make decisions – does not influence the range of options an individual may pick, nor impair their capacity to choose freely. Moreover, the baseline architecture of choice tends to skew decisions toward bad outcomes. Namely, no architecture of choice is neutral, as the rationality of humans is bounded. Therefore, nudging does not impair agency.
{% endhint %}

## Prompt

The evidence previously quoted finally suggest that *choice* overload is different from *information* overload. Namely, the sentiment of discomfort felt by individuals in choice overload does not come from the number of options only, but also from being prompted to choose a subset of options. Intuitively, the disproportion of the option set with the allowed number of picks – too many options, too few picks – makes choosing uncomfortable. This is likely to be especially true in the case of Harambee users selecting skills, as choosing skills to populate a CV is consequential. Being prompted to choose 5 skills is likely to create two issues: 1 – Encourage people to select 5 skills although they are not proficient in some of them. Picking less skills may make users lose confidence by suggesting they have few skills. 2 – Make users feel like they are not able to surface enough of their human capital, which may create frustration.

However, it is unrealistic to populate CV’s with more than 5 skills per occupation. Moreover, the occupation title also sends signals to employers about jobseekers’ skill sets. Therefore, we only recommend adjusting the prompt from “Pick the five skills that you are the most proficient at” to “Pick up to five skills that you are the most proficient at”.

All in all, evidence appear to justify narrowing down the lists of options presented to young jobseekers on the Harambee platform. It also gives evidence that the order of options matters, which may inform decisions taken by Harambee. Finally, it suggests choosing a prompt carefully, so as to not induce stress or demotivation on the platform’s users.

## Bibliography

Neil Malhotra, Completion Time and Response Order Effects in Web Surveys, Public Opinion Quarterly, Volume 72, Issue 5, December 2008, Pages 914–934, <https://doi.org/10.1093/poq/nfn050>.

Reutskaja, E., Iyengar, S., Fasolo, B., & Misuraca, R. (2020). Cognitive and affective consequences of information and choice overload. In *Routledge handbook of bounded rationality* (pp. 625-636). Routledge.

Thaler, R. H., & Sunstein, C. R. (2008). Nudge: Improving decisions about health, wealth, and happiness. Yale University Press.

Wedel, M., & Pieters, R. (Eds.) (2007). Visual Marketing: From attention to action. Hove: Psychology Press.

Xiao, Yaqing and Yan, Hongjun and Davidson, Erik, Response Order Biases in Economic Surveys (February 2, 2024). Available at SSRN: <https://ssrn.com/abstract=3786894> or <http://dx.doi.org/10.2139/ssrn.3786894>.


# Ordering of Skills: A/B testing

As the ordering of skills is likely to greatly influence users' picks, testing different ordering of skill is essential to identify options improving users' experience.

This page summarises a proposal for how to implement A/B testing of how to elicit skills for SAyouth.mobi users’ inclusive CV. This research is motivated by concerns that users may struggle to select between different skills, and an interest in comparing different potential metrics for skill priority. The proposed research is an on-platform randomised evaluation of how users are asked to capture the skills relating to their work experiences.  &#x20;

## Background

In the search to highlight and surface the skills South African youth have as a result of the activities they have performed and are currently involved in, we conducted a survey that asked microentrepreneurs to rate the importance of different skills associated with their occupation. Specifically, we asked about 1500 microentrepreneurs “On the 1 to 5 scale where 1 is ‘Not important at all’ and 5 is ‘Very important’, how important is \[skill tag] to your work?”. The results showed a tendency for participants to assign high importance to all skills presented to them.&#x20;

This poses a potential challenge for the Harambee platform. If job seekers prioritize all skills equally, they might be overly reliant on the initial set of skills they encounter during their platform search. This could lead them to select skills that aren't necessarily the most relevant for their desired careers, hindering their job search effectiveness. To address this concern, our research aims to investigate how the order of skill presentation on the Harambee platform influences job seeker choices.&#x20;

Thus we seek to answer the following questions:&#x20;

1. What is the impact of presenting skills to platform users in different orders on the skills selected?&#x20;
2. Is there an order that offers users the best support in their job search journey?&#x20;
3. Do different skill priorities have downstream impacts on job search patterns?&#x20;
4. Do different skill priorities have an impact on employer interest in the candidate?&#x20;
5. How often should we ask users to update their skills?&#x20;

## Proposal

We propose implementing on-platform A/B testing of how to surface skills to users. This testing should help us to evaluate how users respond to different presentations of skills and consequently refine how best to elicit skill reporting from platform users. &#x20;

The ideal method for testing skills surfacing is evaluating actual user skill selections on platform. This would be achieved by comparing how similar (statistically comparable) users select skills when presented with different list orders and lists of varying numbers of skills. Thus, we wish to randomly assign users to be presented with different lists when they are asked to select the skills they wish to include on their CV. &#x20;

We propose a number of potential skill priority metrics:  &#x20;

1. **Random** \
   This serves as a neutral benchmark. It is also the most theoretically “hands off”.&#x20;
2. **Transferability score** \
   Transferability score = number of appearances of this skill in ESCO / number of occupations in ESCO . This may nudge users towards listing skills that are commonly used in many other opportunities, or in jobs commonly listed on the platform. If users are encouraged to select skills that are more transferable, we may be able to improve the visibility of skills that firms are interested in knowing about. &#x20;
3. **Inverse scarcity of skills** \
   Inverse scarcity of skills  = inverse of number of other users with that score, conditional on the number being above 10 (to rule out idiosyncratic skills). More rare skills may also help job seekers to stand out more from the rest of the field of applicants, and help firms differentiate between job seekers more easily. &#x20;
4. **Some weighted average of these goals** \
   Its likely some mixture of these priorities is appropriate. It may also be that different jobs or different fields may require different weightings of the set of priorities. Developing a preferred weighted average may require some desk work to get input from field experts and literature about relative importance of different measures for particular industries or worker experience profiles. &#x20;
5. **“Other” field to search for skills** \
   Offering users the chance to describe their own skills offers another extreme hands off benchmark. This arm can also be used to validate the accessibility of the language used by the taxonomy to describe the skills surfaced to users.&#x20;
6. **More skill options** \
   For each of the above options we propose surfacing five skills. In this option we propose extending the list to surface more skill options for users to choose from. &#x20;

We will follow a stratified random sampling strategy based on gender and the type of work the individual has performed (i.e. currently unemployed, microentrepreneur, and employed). The sample will be randomly assigned into treatment groups and each group will be presented with a pre-ranked set of skills. The testing could run for one to three weeks depending on logistical constraints for Harambee and the target sample size for each treatment cell. A/B testing in principle does not require detailed power calculations because patterns are informative even if they do not cross the threshold for academic statistically certainty. Samples as small as 20 users per arm could already offer indicative patterns, although on-platform may make it easy to reach far more users. A larger sample would naturally offer the advantages of more precise estimates of effect and the opportunity to account for the role demographic characteristics may play in mediating responses to different skills list orderings. We propose two levels of randomisation.&#x20;

### Across Users Effect

Randomise at the day/week level. While we prefer a day level randomisation, we are happy to work with logistical constraints of bringing various lists live on the platform. We could do every two days, or every week if one of these is easier. Each new randomisation block we will apply a different skill priority ranking. Cycle through each of the treatment arms repeatedly over 3 weeks. This will generate a treatment arm for each list priority order and therefore enable the central comparison of different potential skill priorities.  &#x20;

### Within Users Effect

Once a month prompt an opportunity to review skills  reported  previously and surface skills in a randomly chosen priority (may be the same or different priority users have previously been asked). This will enable us to examine whether the same user reacts differently to receiving skills in different orders, and for users who are shown the same list twice, how stable elicitation is. A measure of skill reporting stability is useful because it can be used to decide how often the platform should prompt users to update their skills or CV. There are a number of reasons to believe that this will be less useful information than the primary randomisation, however, it will be important that the platform has the technical capacity to encourage users to modify which skills they are capturing in their inclusive CV so that we can implement the findings of the A/B testing for platform users in the control group. &#x20;

## Measurement

Many of the indicators of which list to prefer can be collected from administrative data already on the platform. The primary measure of impact is the simple comparison of skills chosen under different treatment conditions. Statistical properties of the skills chosen offer information a number of dimensions of choice quality:&#x20;

1. Correlation between skills chosen under different list orders can tell us the degree to which the different orders make a difference to the skills ultimately chosen.&#x20;
2. Variation in skills chosen within a single priority type can tell us whether all individuals are behaving the same way, indicative of mechanical clicking/less engagement with individual matches with skills or thoughtful engagement.&#x20;
3. Measures of how often the first three skills, or the first skill, or the last skill, or the skill most centrally displayed on the screen can capture indicators or mechanical clicking, or reduced engagement. &#x20;
4. We could manually review a subset of respondents to evaluate skill choice against their profile. &#x20;
5. UX measures of time to complete fields,  number who start and abandon part way through.&#x20;

Demographics which may mediate how users interact with skill lists can also be drawn from user profiles. We would likely focus on:&#x20;

* Gender&#x20;
* Date of birth (Month and Year)&#x20;
* Occupation/activity  &#x20;
* Education level&#x20;
* Geographic location&#x20;
* Experience history (previous years employed)&#x20;
* Disability status&#x20;

All of this information is available for users on the platform, so it would simply be a question of pulling the relevant administrative data.&#x20;

If there is interest in using downstream outcomes to adjudicate between list priorities, we could also draw data on which jobs individuals who have been exposed to different lists click on and apply for. We could also ask firms to select skills they would like to see prioritized among platform applicants from similarly ordered lists to obtain a measure of firm preferences.&#x20;

We could also collect additional data either using on-platform prompts, off-platform follow up surveys, or some small focus groups. Additional data would enable us to measure user satisfaction with the different lists, elicit user criteria for how they select between skills and evaluate how confident or committed users are about the skills they selected. &#x20;

&#x20;


# Evaluating Compass's Performance

{% hint style="danger" %}
Tabiya is working closely with Google and Makesense to develop Compass. Our Minimum Viable Product (MVP) will be tested on real users thanks to Harambee.&#x20;
{% endhint %}


# Takeaways Regarding ESCO

Beyond providing a localised technical answer to Harambee, Tabiya's work in South Africa also allowed us to evaluate the possibility to use the ESCO taxonomy in other national set ups.

Although it is constently improved to include more occupations and skills, the ESCO classification does not yet - and will probably never - cover the full universe of skills and occupations in Europe and, *a fortiori,* in the world. Working with Harambee and working to build an inclusive taxonomy for the South African youth thus called for solutions to adapt and supplement ESCO. These challenges provided us with valuable insights on reproducing our methodology in other countries.&#x20;

{% hint style="info" %}
**How was the ESCO classification built?**&#x20;

The content of ESCO has been developed using "**ESCO reference groups**" for each subcategory of occupations. These groups "\[brought] together experts from different economic sectors and include\[d] employers, education and training providers, trade union representatives, job recruiters, and sector skills council members". An additional **"cross-sector" reference group** also had the duty to develop a vocabulary for transversal skills and competences, which were later used for transversal skills of ESCO occupations. For additional information on the creation of ESCO, refer to this [**paper**](https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=\&arnumber=6928765).  Importantly, ESCO is **not built from scratch**, as it makes use of the EURES taxonomy and existing "supporting" classifications.&#x20;
{% endhint %}

## Internal Inconsistencies in ESCO&#x20;

One of the main takeaways of our work for Harambee is already that ESCO presents some internal inconsistencies. The ESCO classification results from the work of experts, as described above. This human based methodology has advantages:&#x20;

* It allows to avert a data-mining (CVs, job offers) exercise based on data of unequal quality. Job offers and CVs also notoriously overestimate the skills associated with a given occupation.
* Relying on data-mining would have resulted in a taxonomy highlighting the skills that are frequently sought for by employers and displayed by current workers. Therefore, it would tend to invisibilize rare associations of skills and occupations.&#x20;

However, relying on focus groups also creates consistency issues. For instance, the ESCO research group has highlighed the existence of [duplicated skills in the taxonomy](https://www.youtube.com/watch?v=YPECXVRagu8\&t=352s). Because there are multiple reference groups, some conceptually similar skills were created for different occupations under different names. More broadly, we faced two issues. First, the decisions rules used to assign skills to occupations are unclear. Second, using a taxonomy like ESCO entails decisions about a tradeoff between the representativity and the inclusivity of the taxonomy.&#x20;

{% hint style="danger" %}
**Unclear decision rules:** some anecdotal examples highlight internal inconsistencies in ESCO, that are likely due to unclear decision rules. For instance, "nanny" (5311.1.4) is associated with various skills, among which "cook pastry products", whereas "babysitter" (5311.1.2) also requires to know how to cook various products, but not pastries. There is no *a priori* reason why that should be the case, as the two occupations are conceptually very close.&#x20;

This example reveales that the reliance of ESCO on reference groups allows for non-systematic associations of skills with occupations. Although reference groups surely agreed on general rules to decide wether to associate a skill to an occupation or not, this still leaves room for interpretation. Although such inconsistencies do not risk to prevent users of the Harambee platform from picking skills that they have -  platform users always have the option to add skills that they think were missed - **it risks influencing their pick of skills by modifying the architecture of their choices**.&#x20;
{% endhint %}

{% hint style="danger" %}
**Tradeoff between representativity and inclusivity:** Taxonomies like ESCO are meant to impose structure on data (occupations\*skills bundles) that is extremely rich and diverse. Although it is not built on measures of the frequency of the co-appearance of each occupation-skill bundles, relying on human decisions and focus groups still reproduces similar issues: rare occupation-skill associations are necessarily overlooked, to avoid the list of skills associated to occupations becoming too long and, therefore, noninformative and useless.&#x20;

To maximize the representativity and the inclusivity of our taxonomy, while limiting the length of the lists of skills associated witheach occupation, the solution we chose is to keep relying on a "lean" version of ESCO, while allowing platform users to suggest other pre-existing skills when they are assigned a list of skills. When the Harambee platform will go live, we will be able to highlight systamatic occupation-skill associations that are not already included in ESCO, and to potentially decide to enrich our taxonomy.&#x20;
{% endhint %}

## Do we need to extend ESCO?&#x20;

The ESCO taxonomy under its current appears to cover most of the "seen" occupations found in the South African labour market, and the same can be said of skills. The data collected among microentrepreneurs suggests that the latter rarely claim to have occupations or skills that are not already covered by ESCO. Moreover, the Harambee platform aims at offering an "other occupation" and "other skill" option. Therefore, collecting data directly from a large number of job-seekers will allow to evaluate the overall representativity of ESCO in the South African labour market.&#x20;

However, our work with Harambee in South-Africa already convinced us that extending ESCO to include the unseen economy should be at the core of our work, as it allows to build a more inclusive taxonomy. Collecting data about users usage of this part of the taxonomy will allow us to measure the importance of this tool to cover the full range of human capital in the South African labour market.&#x20;

## Adequacy of ESCO for non-European Economies

When it comes seen activities, the microentrepreneurship survey suggests that ESCO satisfactorily covers skills and occupations found in the South African labour market.&#x20;

{% hint style="warning" %}
The data analysis of the microentrepreneurship survey is still ongoing. Final results will be provided hereby shortly.&#x20;
{% endhint %}

As described [here](/tabiya-south-africa/seen-economy/alternative-titles), the main challenge to adapt the inclusive taxonomy to the South African context was to ensure that platform users could access skill and occupation titles that resonate with them. Our experience in South Africa has shown that including alternative titles in our taxonomy was necessary for users to navigate the platform. &#x20;


# Overview

Co-creating open labor market infrastructure with Ethiopia's government

In Ethiopia, Tabiya works with the Ministry of Labor and Skills (MoLS), the Entrepreneurship Development Institute (EDI), and the World Bank to build open-source, AI-powered modules that power the country's national public employment service: a network of Job Centers and a national platform for Ethiopia's labor market information called the E-LMIS.

### Why Ethiopia?

More than two million young Ethiopians enter the labor market every year, and over 70 percent of the workforce is in the informal sector. Urban youth unemployment reached [27.2 percent in 2022](https://ess.gov.et/wp-content/uploads/2024/09/2022_-2nd-round-UEUS-Key-Findings.pdf), and the rate for young women was roughly twice that for young men. Where opportunities are this scarce, it matters all the more that the few available reach the people who need them, yet the labor market is fragmented. Job postings are unstructured free text that algorithms cannot read, there is little real-time visibility into where skills are needed, and compliance checks are done by hand across thousands of centers. The people with the fewest connections are the least likely to hear about an opening in time; urban youth have applications on their devices while rural citizens travel for hours, and jobseekers looking for work abroad are exposed to unregulated brokers.\
\
At the same time, Ethiopia is among Africa's fastest-growing economies: the IMF's April 2026 outlook projects growth of around 9 percent in 2026. But aggregate growth, led by sectors such as gold, coffee, and aviation, does not automatically reach the jobseekers who are hardest to place, among them women, less-educated jobseekers, migrants, and refugees. That gap between headline growth and individual livelihoods is a large part of the case for stronger career guidance and intermediation.

A public employment service, delivered locally and embedded in government, can reach jobseekers that online platforms tend to miss, which makes strengthening it a matter of inclusion as much as efficiency. This is the gap Ethiopia's Job Centers and the E-LMIS exist to close, and where Tabiya's tools are applied.

### Our Partners

The [Urban Productive Safety Net and Jobs Project](https://documents.worldbank.org/en/publication/documents-reports/documentdetail/099031324145511013) (UPSNJP), financed by the World Bank and led by the Ministry of Urban Development and Infrastructure, supports the urban poor and the labor market inclusion of disadvantaged urban youth across 88 cities, reaching roughly 1.7 million people.\
\
The [E-LMIS](https://lmis.gov.et/) led by the Ministry of Labor and Skills is built around three goals: a competent employee matched to the right opportunity, a competent employer receiving curated and qualified shortlists, and a competent decision-maker with real-time policy data. The [Entrepreneurship Development Institute ](https://edi-ethiopia.org/)(EDI), a quasi-governmental institute accountable to MoLS supports Job Center capacity building and training. The World Bank's UPSNJP convenes and funds the wider project.

{% columns %}
{% column width="33.33333333333333%" valign="middle" %}

<figure><img src="/files/9Vk6bq9BQNm4q3RjCKq2" alt="Entrepreneurship Development Institute logo"><figcaption></figcaption></figure>
{% endcolumn %}

{% column width="33.33333333333333%" %}

<figure><img src="/files/beJIwhIewzwppsUirjen" alt="World Bank logo"><figcaption></figcaption></figure>
{% endcolumn %}

{% column width="33.33333333333335%" valign="middle" %}

<figure><img src="/files/i4a2eaTa2wPzGeM7eklX" alt="Entrepreneurship Development Institute logo" width="441"><figcaption></figcaption></figure>
{% endcolumn %}
{% endcolumns %}

<figure><img src="/files/mcEeau6lWSbKHrXscin9" alt="" width="375"><figcaption></figcaption></figure>

This system already reaches significant scale: more than four million registered job seekers and some 13.7 million job orders, running through 787 job centers within a national network of 2,452 woreda job centers. Tabiya is the technology and thought partner, contributing the taxonomy and matching engine at the heart of the platform.


# Localized Taxonomy

Localization decisions, mappings, and data.

A job seeker in Addis Ababa might describe their experience in Amharic, in informal and colloquial terms, drawing on work that never carried a formal title. On the other hand, employers may post a vacancy that potentially describes similar work but with a different vocabulary. Good labor-market matching has to recognize the full universe of skills that work can be described in. Building on the experience in South Africa and Kenya, Tabiya supported the E-LMIS team at MoLS to map noisy, inconsistent occupation labels onto a clean, standardized occupational framework for Ethiopia, aligned with international standards through localization of the Inclusive Livelihoods Taxonomy .&#x20;

The outcome of this effort is a “gold standard” of roughly 9,000 canonical occupational positions distilled from more than 130,000 raw historical entries, drawn from ESCO, ISCO, ILO, and the Ethiopian Civil Service, each mapped to skill tags and competency vectors. The sections below walk through how that was done, from raw data to a finished catalog.

<details>

<summary><strong>What did Tabiya and the E-LMIS team set out to build?</strong></summary>

The E-LMIS team compiled a dataset of raw occupations sourced from several sources that it wanted to anchor the national Ethiopian taxonomy on. The teams processed these raw vacancies into a clean, standardized catalog through a multistep taxonomy localization process. The main tool used for this effort was the Tabiya's [Livelihoods Classifier,](/our-tech-stack/livelihoods-classifier) which reads free-text occupation labels and maps them to standardized occupations and skills on Tabiya's ESCO-based Inclusive Livelihoods Taxonomy.&#x20;

</details>

<details>

<summary><strong>Step 1: Sourcing the vacancy dataset for taxonomy localization</strong></summary>

The work drew on two separate streams, kept apart during collection so that operational, user-generated data would not contaminate the authoritative reference data.&#x20;

1. **E-LMIS operational dataset:** roughly 131,000 raw occupation records that users had entered through different workflows, containing duplicate entries, typographical inconsistencies, departmental and company-specific names, experience-level prefixes, and multiple naming variations for the same occupation. An initial round of cleaning and deduplication reduced these to about 9,000 unique occupation labels.&#x20;
2. **Reference dataset of \~9,700 records from authoritative sources**: roughly 700 ISCO-aligned global occupations from the International Labour Organization (ILO), about 1,000 occupational descriptors and skill structures from the O\*NET Resource Center, and around 8,000 localized public-sector occupation structures from the Ethiopian Civil Service.

The two streams were merged into an intermediate pool of about 18,000 records. A further pass removed duplicate titles, structural overlaps (the same role written, ordered, or phrased differently), invalid records, and semantically redundant labels leaving roughly 9,000 cleaned and standardized occupation titles to use for taxonomy alignment and recommendation processing.

</details>

<details>

<summary><strong>Step 2: Pre-processing and cleaning the vacancy dataset before classification</strong></summary>

Before alignment, a Large Language Model (LLM)-assisted normalization layer was introduced to rewrite noisy, user-entered titles into the core occupational concept they represented. This step removed organization and workplace names (for example, “Administrator at Terra LAB” became “Administrator”), stripped spoken-language descriptors (“Afan Oromo and Tigrigna Translator” became “Translator”), consolidated variations of the same role (“Administrative Assistant Officer” became “Admin Assistant”), separated titles that bundled two jobs into one (“Accountants and Auditors” became separate “Accountant” and “Auditor” entries), and removed seniority or experience prefixes (“Junior Accountant” became “Accountant”). The resulting dataset trimmed noise and retained consistent occupation names for similar clusters of work.&#x20;

</details>

<details>

<summary><strong>Step 3: Mapping locally derived occupation titles to a standard taxonomy</strong> </summary>

The cleaned titles were first mapped using the Classifier on to the Tabiya (ESCO 1.1.1) v2.0.1-rc.1 taxonomy base. A named entity linking (NEL) step matched each cleaned title to a taxonomy occupation using semantic similariy and assigning a match only when confidence passed a 0.80 similarity threshold. The threshold was selected after several rounds of iterative human review of Classifier outputs. This step left about 2,000 low-similarity titles below the cutoff. For high similarity, non-exact matches, validated local occupation titles were added to the taxonomy as alternative labels, improving occupational searchability and localization.&#x20;

*Note: Amharic labels were excluded from this synchronization cycle to keep formatting consistent across taxonomy structures.*

</details>

<details>

<summary><strong>Step 4: Human review and expanding the taxonomy - pending post-pilot</strong></summary>

As of August 2026, the taxonomy localization process is pending a targeted manual review of the roughly 2,000 unmatched/low match score titles, ensuring full dataset coverage. This manual review process will include additions to the occupations' repository beyond simply mapping to existing occupations, and explore methodologies to validate the skills associated with localized Ethiopian occupations.&#x20;

</details>

The taxonomy is the foundational layer that will power the end-to-end jobseeker and vacancy pipeline on the national employment portal. It currently standardizes Ethiopia's labor market data at the source and powers skills-based matching, so that the resulting data can add up to a timely picture of labor supply and demand that the government can use for planning and policy.


# Skills-matching algorithm

Platform, partners, and integrations.

Ethiopia's public employment service is delivered through a network of physical Job Centers managed by MoLS and its digital backbone, the E-LMIS . Tabiya supports the redesign of the Job Centers' workflows and operating procedures and has shared its open-source matching engine with the E-LMIS team, whose engineers are integrating it into their new platform.&#x20;

### Integrating an improved skills-based matching algorithm

Previously, the E-LMIS matched jobseekers to vacancies with a basic, deterministic rule based on skill overlap, and its skills data was a flat list of \~25,000 entries not aligned to any international standard. Tabiya's  matching service replaces this rule-based matching with a skill-to-skill matching algorithm that compares every skill a jobseeker holds against every skill a vacancy requires in a shared semantic space, so that closely related skills, not only identical ones, count toward a match. Each required skill is scored against the jobseeker's nearest equivalent on a continuous 0–1 scale.  Skill similarity is the core signal, which the service then combines with location, jobseeker preferences, education eligibility, and labour-market demand to produce a final ranking.

***

### *How the matching engine works*

Matching runs on a structured skills profile, not keyword search. The engine compares what a person can do against what a vacancy requires, weighs that against what the person wants and how likely an employer is to hire them, and returns a ranked shortlist a counsellor and jobseeker can act on.

{% stepper %}
{% step %}
***Map both sides onto a shared skills taxonomy***

*Jobseeker profiles and vacancy descriptions are resolved onto the same \~14,000 skills repository of the localized Ethiopian taxonomy, so a skill* *described differently on a CV and in a job advert resolves to a single entry.*
{% endstep %}

{% step %}
***Place every skill in a semantic space***

*Each skill is represented as a vector positioned by meaning, letting the engine measure how close any two skills are rather than only whether they are identical. This is what allows related and transferable skills to count toward a match.*
{% endstep %}

{% step %}
***Shortlist the closest candidates with a fast pass***

*Each jobseeker's skills are averaged into one vector, and that profile is compared against the job in one fast sweep. Using FAISS, the candidates whose skills sit closest to what the role needs rise to the top; the nearest 50 or so continue.*
{% endstep %}

{% step %}
***Score each shortlisted candidate skill by skill***

*For every skill a vacancy requires, the engine finds the jobseeker's closest skill and scores that pair from 0 to 1. Essential requirements are combined strictly, so a serious gap in a must-have skill pulls the score down and cannot* *be offset by strength elsewhere. Optional skills add credit more gently.*
{% endstep %}

{% step %}
***One final ranking***

*The skill score is combined with the jobseeker's stated preferences and with an estimate of how likely an employer* *is to hire them — essential-skill coverage, wider readiness, and current demand for the role.*&#x20;
{% endstep %}
{% endstepper %}

### &#x20;*Running on models the government has full ownership of:*&#x20;

The matching service depends on models to classify vacancies, place skills in semantic space and to re-read the shortlist. These can be commercial services reached over an API, or open-weight models hosted on infrastructure the government runs. As part of this work, the two teams are testing and optimizing the engines for the second option.

Local deployment keeps jobseekers' personal data on systems the government administers rather than passing it to an external provider, addressing data sovereignty requirements for a national system holding personal records. It also removes the per-request cost of matching, which makes the service more affordable to operate at the volumes a full Job Center network generates.


# Roadmap

### *Where things currently stand*

The skills-based matching approach Tabiya developed with government partners in South Africa and Kenya now runs on the E-LMIS team's own infrastructure. The team has deployed the matching service locally, generating skill embeddings with an open-weight model on their own servers, using a localised Ethiopian taxonomy (refinement in progress) of roughly 14,000 ESCO-aligned skills.

Tabiya serves as the ministry's main technology and technical-assistance partner for the E-LMIS, co-leading the digitalisation workstreams. The two teams have localised the base taxonomy — around 80% complete, with human review  continuing on the remainder and it is already in use in the matching pilot. Tabiya has also helped the E-LMIS team define key performance indicators and validate the metrics being built into their analytics platform. A taxonomy management application, which would make the taxonomy a maintained living resource rather than a static dataset, is being scoped.

***Piloting**: The collaboration positions the 18 model Job Centers as a demonstration site for how open-source digital public goods and AI-enabled tools can strengthen government-led employment services. A guiding principle is that these tools support rather than replace human-delivered services: their outputs are decision-support for counselors, not automated determinations of eligibility, benefits, or employment outcomes.*&#x20;

### *What's next?*

MoLS and the E-LMIS team have scoped around nine priorities for the modernisation of Ethiopia's public employment system. Tabiya is an active delivery partner on four of them - the living Ethiopian taxonomy, skills extraction and vacancy structuring, skills-based matching, and labor intelligence layer, and provides continued advisory and technical guidance across the rest.

These four form a single chain. The **taxonomy** gives every occupation and skill a shared name. **Extraction** turns unstructured vacancy text into structured skills expressed in that vocabulary. **Matching** compares supply and demand once both sides are anchored in it. An **analytics and intelligence layer** opens the system's growing supply- and demand-side data to non-technical decision-makers through plain-language querying, and enrich it with signals from elsewhere in government.

**Inclusive and bias-aware career guidance** is an upcoming priority. It determines whether a jobseeker's skills including those built in informal, unpaid and household work — enter the system accurately in the first place. Eliciting someone's skills and helping them think about their working future are the same conversation from the jobseeker's side, and the E-LMIS  team's plans for its My Career and MyFuture portals overlap directly with Tabiya's Compass, deployed in Argentina, Kenya, South Africa and Zambia.

Across the remaining priorities, Tabiya's role is advisory: helping scope what is feasible, and where its open-source components could contribute. Each depends on funding, data access, and decisions that sit with the ministry. These include a National Recruitment Platform, verifiable credentials built on Ethiopia's Fayda digital ID, the packaging of E-LMIS components as digital  public goods, and a national skill bank and skill libraries. Each depends on funding, data access, and decisions that sit with the ministry.


