🇨🇴⚖️ La Rama Judicial valida a Ariel en prueba de concepto de IA. Conoce los resultados aquí

OIT - Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques

OIT - Organización Internacional del Trabajo

Icono de documento PDF

Descargar PDF

Disponible

Detalles

Título
OIT - Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques
Autor
OIT - Organización Internacional del Trabajo
Categoría
Doctrina
Área del derecho
Laboral
Año

 ILO Brief 1 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques  Methodological Brief January 2025 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques Willian Adamczyk†, Simon Boehmer†, Isaure Delaporte†, Verónica Escudero†, Hannah Liepmann†

 This methodological brief describes an innovative approach that exploits online big data from job vacancies and applicants' profiles and utilizes natural language processing (NLP) to extract information on skills.  The approach is based on a taxonomy comprising 15 unique skills subcategories across the broader cognitive, socio -emotional, and manual skills categories. It focuses on skills that are transferable rather than occupation-specific.  So far tested with data on Uruguay, Brazil, the Russian Federation, and South Africa, t he method is applicable to countries of all income levels and allows detecting country-specific skills trends. It captures a wide range of sectors and occupations, including those requiring manual labour.  It can be used to offer novel insights into skills demand, supply and mismatch , as well as into the relationship between skills and aspects of job quality and/or job transitions.  Introduction As global labour markets undergo rapid transformation , understanding the dynamics of skills is crucial for informed policy-making and economic development. In Western contexts, considerable scholarly efforts have been directed towards employing skills classifications to discern overarching skill trends and examine how skills dynamics influence wages and employment (see Autor, Levy, and Murnane 2003 ; Acemoglu and Autor 2011 ; Deming and Kahn 2018; Atalay et al. 2020; Hanushek et al. 2017). These studies have yielded critical insights into the influence of skill formation on workforce productivity, income

Murnane 2003 ; Acemoglu and Autor 2011 ; Deming and Kahn 2018; Atalay et al. 2020; Hanushek et al. 2017). These studies have yielded critical insights into the influence of skill formation on workforce productivity, income distribution, and economic resilience. However, in contrast

We would like to thank Evgeny Gushchin, Elvire Jégu and Franziska Riepl who provided research support for this brief. † Skills, Active Labour Market Policies and Policy Evaluation Team at the Research Department of the ILO. to the substantial focus on Europe and especially the United States, the exploration of similar inquiries remains relatively underdeveloped in emerging economies. The existing literature draws upon diverse sources of information on skills. One prominent approach has been to use detailed occupational classification systems such as the Occupational Information Network (O -NET) of the United States and the European Skills, Competences, Qualifications and Occupations ( ESCO) classification . However, the transferability of these occupational classifications to different national contexts is challenging , Key points ILO Brief 2 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques given that the skills composition within occupations differs significantly across countries. In addition, O-NET and ESCO represent extremely granular ontologies (i.e., comprehensive listings of relevant skills) that require conceptual aggregation before they can be used in research. Another approach utilizes surveys designed to measure skills, such as the OECD’s Programme for the International Assessment of Adult Competencies ( PIAAC) and the World Bank’s Skills Toward Employment and Productivity (STEP) surveys (OECD 2019; World Bank 2014). These initiatives offer valuable insights. However, their limited coverage in terms of countries and years —often a single year per country —constrains their applicability for achieving a comprehensive global perspective . Given the dynamic nature of tasks and skills in response to evolving

These initiatives offer valuable insights. However, their limited coverage in terms of countries and years —often a single year per country —constrains their applicability for achieving a comprehensive global perspective . Given the dynamic nature of tasks and skills in response to evolving labour market s, there is a need to longitudinally assess skills within and across occupations as well as in various country contexts. This methodological brief presents an innovative approach which was first developed in Escudero, Liepmann, and Podjanin (2024) and applied on Uruguayan data, and then further refined and extended to additional countries — Brazil, the Russian Federation, and South Africa —for the forthcoming 2026 World Employment and Social Outlook (WESO) Report on Lifelong Learning and Skills Dynamics . This method leverages big online data and advanced natural language processing (NLP) techniques to reveal skills trends in low - and middle -income countries . The purpose of this brief is hence to provide a detailed description of the methodological fram ework, including data processing and analysis steps . It is meant not only to enable the replication of analys es featured in the WESO Thematic report and accompanying research projects, but also to support adaptation of the methodology for further research on skills trends in countries of different income levels. This brief thus aims to serve as a practical guide for researchers and practitioners seeking to apply these techniques in their own studies. The use of o nline data on labour markets has in recent years increased substantially (see the discussion in Fabo and Kureková 2022 ). The methodology presented in this brief offers several distinct advantages. First, it leverages the nature and granularity of the data . By extracting detailed information from job vacancies and applicants’ profiles, this approach addresses significant data gaps, enabling country-specific analyses without assuming occupational skills are similar across countries . It also allows to measure skills over time as they are presumably evolving in response to labour market transformations . Second, the proposed taxonomy is particularly adaptable

profiles, this approach addresses significant data gaps, enabling country-specific analyses without assuming occupational skills are similar across countries . It also allows to measure skills over time as they are presumably evolving in response to labour market transformations . Second, the proposed taxonomy is particularly adaptable to the realities of lowand middle-income countries. Unlike prior skills taxonomies for online vacancy data, which primarily focus on professional jobs in high -income countries (see, for example, the seminal study of Deming and Kahn 2018 ), this taxonomy incorporates manual skills that are relevant across countries, but even more so in lowand middle -income economies . It also expands the conceptual foundation of socio -emotional skills. Third, the taxonomy prioritizes transferable skills —those applicable across occupations —over occupation-specific skills , while highlighting numerous technical skills, making it particularly useful for understanding broad labour market trends. In summary , this methodology provides a robust framework for analysing labour market dynamics . This in turn facilitates targeted interventions and policy formulation to enhance workforce development and economic growth. The remainder of this brief presents the key elements of the methodology applied to create skills variables. It also discusses how, in a similar way, open-text descriptions can be used to extract information and create an ISCO -08 occupation variable. The methodology has already been applied to answer pertinent skills-related questions in a number of studies within the ILO (see Escudero and Riepl 2024 using South African and Uruguayan data; De Marzo, Mathew, and Sbardella 2023 using Indian data ; and Escudero, Liepmann, and Vergara 2024 using Uruguayan data).  Methodology to create skills variables The methodology encompasses skills related to both jobspecific tasks and personal attributes. This inclusive approach ensures a comprehensive representation of skills-related dynamics across diverse labour markets , covering the skills sought by employers in job postings and those highlighted by workers in their online profiles. In the

skills variables The methodology encompasses skills related to both jobspecific tasks and personal attributes. This inclusive approach ensures a comprehensive representation of skills-related dynamics across diverse labour markets , covering the skills sought by employers in job postings and those highlighted by workers in their online profiles. In the following, the basic building blocks of the original taxonomy—developed in Escudero, Liepmann, and Podjanin (2024)—are presented along with the refinements that were implemented in further work. ILO Brief 3 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques Taxonomy In the taxonomy, skills are grouped into three broad categories, namely cognitive, socio -emotional and manual skills. Within each category, further distinct skills are identified to arrive to a classification comprising 15 subcategories (14 in the original concept) , making it a nuanced yet compact taxonomy. The taxonomy is built upon existing literature from labour economics and psychology , but has been expanded to make it adaptable to individual country -contexts, with a particular focus on emerging and developing countries and applications in online data. Initially, the skills categorization was developed based on established taxonomies designed for classifying skills in online data within the United States (U.S.), particularly Deming and Kahn (2018).1 The taxonomy was then expanded to include manual skills, which are typically omitted in U.S. -centred analyses of online data . The taxonomy includes three skills subcategories for manual skills (i.e. finger dexterity, hand -foot-eye coordination, and physical skills) to facilitate a more comprehensive analysis of online data beyond individuals with high formal qualifications . Moreover, the conceptual foundations of cognitive and socio -emotional skills were broadened by integrating additional keywords that capture shifts in terminology over time while still represe nting the same skill (Deming and Noray 2020). To achieve this, a dictionary of keywords and expressions referring to each skill subcategory was developed. These

broadened by integrating additional keywords that capture shifts in terminology over time while still represe nting the same skill (Deming and Noray 2020). To achieve this, a dictionary of keywords and expressions referring to each skill subcategory was developed. These keywords and expressions were drawn from various seminal studies, some of which do not rely on online data sources (Almlund et al. 2011; Atalay et al. 2020; Deming and Noray 2020; Heckman and Kautz 2012; Hershbein and Kahn 2018; Kureková et al. 2016; Spitz‐Oener 2006), as well as the pilot version of O-NET Uruguay.2

As mentioned, further improvements were made to the initial version of the taxonomy. The changes include:

1. Splitting the cogni tive skills subcategory into core cognitive and sophisticated cognitive skills. The updated version of the taxonomy thus includes 15 subcategories of skills.

2. Certain skills subcategories have also been refined: i) the former project management skills subcategory has been renamed to project and process management skills and now comprises additional types of skills; ii) as a result, the people management skills subcategory has been refined and now refers to skills only specific to the management of people.

3. Lastly, the list of keywords and expressions has been updated following a new round of context revision. For instance, to account for modern technologies, programming language and software, additional terms were retrieved from topics in Stack Overflow and Github, the two most popular platforms amongst software developers, and were added to the list of keywords. Similarly, expressions related to machine learning and artificial intelligence were added to the list of keywords using relevant tagged questions in Stack

Overflow. Table 1 presents an overview of the skills categories and subcategories, along with their definitions and sources they were derived from . Additionally, Table A1 in the Appendix lists the most important identifying keywords

of keywords using relevant tagged questions in Stack Overflow. Table 1 presents an overview of the skills categories and subcategories, along with their definitions and sources they were derived from . Additionally, Table A1 in the Appendix lists the most important identifying keywords and details the changes made compared to the original version presented in Escudero, Liepmann, and Podjanin (2024).3 For more details, researchers and practitioners seeking to replicate the method are encouraged to contact the ILO through the contact details provided at the end of this brief.

1 Other sources used include Deming and Noray (2020); Heckman and Kautz (2012) and Kureková et al. (2016). 2 O-NET Uruguay is an occupational classification system which describes the tasks and skills associated with a specific occupation in the Uruguayan context, similar to the US O-NET. When the methodology of this brief was developed, the O-NET Uruguay pilot project captured 22 selected occupations only (see Ministerio de Trabajo y Seguridad Social 2020; Velardez 2021). 3 The complete dictionary for each language will be made publicly available as part of the work undertaken for the forthcoming 2026 World Employment and Social Outlook (WESO) Report on Lifelong Learning and Skills Dynamics . ILO Brief 4 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques  Table 1: Categorization of skills, keywords, and sources Subcategory Definition Sources Cognitive skills Cognitive skills (core) Skills needed to perform tasks that require analysis and calculation, problemsolving, intuition, and flexibility. DK (2018); ALM (2003); DN (2020); S-O (2006) Cognitive skills (sophisticated) Skills needed to perform more sophisticated tasks that require analysis, modelling, and creativity.

DK (2018); ALM (2003); DN (2020); S-O (2006)

(2020); S-O (2006) Cognitive skills (sophisticated) Skills needed to perform more sophisticated tasks that require analysis, modelling, and creativity. DK (2018); ALM (2003); DN (2020); S-O (2006) General computer skills These six subcategories relate closely to the cognitive skills described above. They correspond to skills that are needed in specif ic areas of work. They are listed as separate subcategories because they are often specifically mentioned in job adverts and in applicants’ work experiences. They are mostly geared towards white-collar jobs, in line with the aim of the work by Deming and Kahn (2018).

APST (2020); O-NET Uruguay; DK (2018); DN (2020) Software skills and technical support Machine learning and AI Financial skills Writing skills Project and process management Socio-emotional skills Character skills Character skills include three of the five categories of the five -factor personality model commonly used in the psychological literature (McCrae and Costa 2008). It includes conscientiousness, openness to experience, and emotional stability. In addition, this subcategory includes dimensions such as being relaxed, independent, self-confident and the degree of vulnerability to stress. DK (2018); DN (2020); KBHT (2016); HK (2012) Social skills Social skills include those character traits from the five -factor personality model that are less related to one’s personal attributes and related more to how one interacts with other people, specifically agreeableness and extraversion. Other keywords that relate to the general ability to have personal interactions, such as working in teams or holding presentations, are included as well. DK (2018); DN (2020); KBHT (2016); HK (2012); S-O (2006); APST (2020); O-NET Uruguay People management skills Lastly, two subcategories are added that refer to specific abilities within the

as well. DK (2018); DN (2020); KBHT (2016); HK (2012); S-O (2006); APST (2020); O-NET Uruguay People management skills Lastly, two subcategories are added that refer to specific abilities within the broader realm of social interactions, and which are often listed as particular requirements in job adverts and applications. DK (2018); DN (2020); Customer service skills ALM (2003); S-O (2006) Manual skills Finger-dexterity skills This category focuses on manual skills that are usually classified as “routine” by ALM (2003) and which are common in machine operation and the production or handling of goods. Examples include picking and sorting in agriculture or working in an assembly line.

ALM (2003); S-O (2006);

APST (2020); O-NET Uruguay Hand-foot-eye coordination skills These manual skills are usually understood as “non -routine” by ALM (2003). They are more commonly used in services -related occupations and include working in changing environments that necessitate adaptation. This category encompasses for example driving cars or repairing and cleaning items. ALM (2003); S-O (2006); PST (2020); O-NET Uruguay Physical skills This subcategory focuses on more innate bodily characteristics, such as physical strength, endurance, the ability to lift heavy objects or work while standing or walking.

O-NET Uruguay Source: Table 1 of Escudero, Liepmann, and Podjanin (2024), where additional information on the conceptual background and concrete keywords used in the taxonomy can be found. The latest list of keywords and expressions can be found in Table A1 of the Appendix.

Notes: ALM (2003) stands for Autor et al. (2003), APST (2020) for Atalay et al. (2020), DK (2018) for Deming and Kahn (2018), DN (2020) for Deming and Noray (2020) , HK (2018) for Hershbein and Kahn (2018) , HK (2012) for Heckman and Kautz (2012) , KBHT (2016) for Kureková et al. (2016) and S-O (2006) for Spitz-Oener (2006). To keep pace with changing market dynamics, the skills taxonomy and associated dictionaries should be regularly updated and revised. ILO Brief 5

Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques Implementation This taxonomy is then applied to the online vacancy and applicants’ data combining rule -based classification— derived from the taxonomy —with natural language processing (NLP) techniques. These techniques allow to systematically access and review unstructured text information contained in big data. More specifically, opentext descriptions of vacancies are pre-processed to fit the structured format of the skills taxonomy and its list of keywords. Skills are identified based on the specific dictionary of keywords and expressions associated with each of the 15 skills subcategories in the taxonomy. Box 1 provides an overview of the data used so far.  Box 1. Data The methodology was applied and adapted to four country contexts: Uruguay, Brazil, the Russian Federation, and South Africa. The data used comes from the following sources:  BuscoJobs, a private job -search portal where firms can post vacancies for a small fee and applicants can create profiles and apply to the vacancies. It provides detailed information on job vacancies posted by firms, applicants searching for jobs, and on the applications made by job seekers to those vacancies in Uruguay and Brazil. The data we use span years 2010 through 2023.

create profiles and apply to the vacancies. It provides detailed information on job vacancies posted by firms, applicants searching for jobs, and on the applications made by job seekers to those vacancies in Uruguay and Brazil. The data we use span years 2010 through 2023.  Adzuna, a job aggregator which collects, standardizes and re -posts vacancies published on the internet. It provides detailed job advert data in the Russian Federation and South Africa . For both countries , we use data covering the period from April 2016 through December 2021. Although this methodology has been implemented and tested specifically in the context of these four countries, the growing availability of online data presents a n opportunity for applying the method s elsewhere. As such, t he method ology holds significant potential for generating insights into country -specific labo ur dynamics across a broader array of global contexts. One such example is the study by De Marzo, Mathew, and Sbardella (2023), which uses vacancy data to investigate the link between skills demand and firm s’ productivity and innovation in India. Although some skills subcategories are closely linked, the keywords and expressions used to characterize them are distinct and mutually exclusive, which allows for the unique identification of skills in the data. When keywords overlap in different categories, an exception rule was added to avoid partial matches. Such cases are for instance: “design” which is included in core cognitive skills and “design site” which is in software and technical support skills; or “repair” in hand-foot-eye coordination skills and “computer” in computer general skills , which do not match “repair computer” in software and technical support skills; or lastly, “team” in social skills, which is treated separately from “lead team” or “team management” in people management skills. To adapt the dictionary to country -specific languages and contexts, keywords were translated while considering local expressions. Some modifications were made to the original keywords to avoid ambiguity or miscontextualization after

team” or “team management” in people management skills. To adapt the dictionary to country -specific languages and contexts, keywords were translated while considering local expressions. Some modifications were made to the original keywords to avoid ambiguity or miscontextualization after they had been translated. To ensure correct usage, rounds of manual inspection were conducted by the authors of this brief to verify whether the suggested word was suitable in at least 75 per cent of descriptions in 40 randomly selected observations. If it was found not to be suitable, whenever feasible another synonym or compound expression was chosen instead of dropping altogether the direct ly translated keyword. For instance, the keyword “running”, as listed in the taxonomy under physical skills, can be translated as “correr” in Spanish and Portuguese dictionaries, which accurately reflects the physical context. In English, however, the word “running” can refer to non -physical tasks as well such as “running a project”. Therefore, this keyword was omitted from the E nglish dictionary. Synonyms like “run(ning) fast” or “athletic (run)” were considered non-useful as they did not appear frequently in job postings or wer e used in incorrect contexts. Still, the physical aspects of “running” were conveyed through terms like “walking”, “strolling”, “hiking”, or “marching”. The dictionaries were developed specifically for each country as follows: • Uruguay: This dictionary was developed and implemented first, based on the translation to Spanish of the English dictionary that was elicited from the literature review in Escudero, Liepmann, and Podjanin (2024 ). After translation, terms that consisted of two words were also added in reverse order and synonyms were added based on the Spanish version of the synonym website www.wordreference.com. The added synonyms were ILO Brief 6 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques manually reviewed to discard out-of-context or overlapping terms. This yielded first a set of 275 keywords and

www.wordreference.com. The added synonyms were ILO Brief 6 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques manually reviewed to discard out-of-context or overlapping terms. This yielded first a set of 275 keywords and expressions4 (without synonyms), which increased to a final set of 669 keywords and expressions after synonyms had been included. • Brazil: the original English dictionary was translated to Portuguese, taking into account the adjustments that were made for the Spanish version. Synonyms from the website www.sinonimos.com.br were added, yielding first a total of 306 keywords and expressions, and finally 1,339 keywords and expressions including synonyms. • The Russian Federation : The English dictionary was translated into Russian taking into account potential differences in context.5 In an additional step, software and company names were also added in the Latin alphabet, and selected keywords with ambiguous meanings were modified or dropped. Synonymys were scraped from https://synonymonline.ru/ and refined manually. The final dictionary contains 303 keywords and expressions and 1,951 such terms when including synonyms. • South Africa: This dictionary was built directly from the English keywords derived from the literature. Synonyms were added by performing an automated web scraping of the synonym website www.wordreference.com. This procedure yielded a final set of 284 keywords and expressions or 1,686 with the inclusion of synonyms. Overall, w eb scraping procedures return a different number of synonyms for each language, owing to its specific linguistic characteristics. The obtained synonyms were reviewed manually to confirm their relevance and to ensure comparability across countries . N evertheless, the different number of synonyms does not substantially affect the number of matched skills for each country. Creation of the skills variables The final step i nvolves creatin g the skills variables by leveraging the unstructured text data found in vacancy descriptions posted by firms and the job spells listed in applicants’ profiles. The open -text descriptions offer a

the number of matched skills for each country. Creation of the skills variables The final step i nvolves creatin g the skills variables by leveraging the unstructured text data found in vacancy descriptions posted by firms and the job spells listed in applicants’ profiles. The open -text descriptions offer a viable approach for creating skills variables, as they contain

4 Keywords refer to single wo rds, whereas expressions denote multiple words belonging together. 5 The contributions of Evgeny Gushchin to the skills variable creation for the Russian Federation are gratefully acknowledged. detailed information on skills for all vacancies and a majority of applicants’ job spells. The open -text descriptions undergo a series of NLP preprocessing steps, including: (i) tokenization (i.e., splitting the text s into their individual words to allow for further processing), (ii) text normalization (e.g., converting to lowercase and removing accents , numbers , special characters, and words with fewer than two letters), (iii) stop words removal (i.e., dropping words that do not carry meaning for the exercise at hand) , and (iv) lemmatization (i.e., reducing words to their common root to unify variations of the same concept, such as “communicate” and “communication”, while accounting for gender, plural, and verb tense variations). These processes are applied to both the online data and the skills taxonomy categorizations, to facilitate the mapping of the two. Finally, each skill subcategory is considered and coded as present if at least one of the keywords from the dictionary is identified in the text. Additionally, the frequency of keyword occurrences for each skill subcategory is calculated6 and used to study the overall supply and demand for skills. The NLP implementation largely follows the same process in all four countries, except when different sets of tools and algorithms were necessary to accommodate data size and language-specific algorithms: • Uruguay: as the total number of observations was

calculated6 and used to study the overall supply and demand for skills. The NLP implementation largely follows the same process in all four countries, except when different sets of tools and algorithms were necessary to accommodate data size and language-specific algorithms: • Uruguay: as the total number of observations was 164,864 vacancies and 1.5 million job spells , the data was processed locally using Python’s Natural Language Toolkit (NLTK) to tokenize, normalize, drop stopwords and lemmatize words. • Brazil: as the data comprises 42 million vacancies and 3.2 million applicants’ job spells, PySpark and SparkNLP were required to tokenize, normalize, drop stopwords and lemmatize words. • The Russian Federation: the data comprises 170 million observations for vacancies ( there is no information on applicants). However, due to performance restrictions, a random sample of 10 per cent of the data was used. With 17 million, the resulting sample is large enough to detect all relevant patterns in the data. Similarly to the case of 6 While this frequency provides valuable insights, it does not fully capture the intensity with which a particular skill is used. This constitutes a caveat that warrants further exploration in future research. ILO Brief 7 Developing a New Method to Uncover Skills Trends in Emerging Economies Using Online Data and NLP techniques Brazil, PySpark and SparkNLP were required to tokenize, normalize, drop stopwords and lemmatize words. • South Africa: the data contains 6.2 million observations (again only of vacancies). This sample size required using the big data tools present in PySpark and SparkNLP. Turning to the classification outcomes, applicants’ descriptions of their past jobs tend to be shorter than vacancy postings for similar jobs, and thus capture less skills on averag

Estás viendo una vista previa

Lee el documento completo con Ariel

Este es un fragmento de uno de los más de 1.2 millones de documentos de la biblioteca de Ariel. Crea tu cuenta para leerlo completo, descargarlo y consultarlo con Ariel, que siempre te lleva a la fuente exacta: Ariel NO alucina.

Consultar sobre este documento ...