Data and methods
Thirteen linked tables on 22 states, open under CC BY 4.0, with the methods used to analyse them.
How to cite
Ben Brik, A., Gilbert, N., & Pycińska, M. (2025). Arab AI Governance Lab Dataset [Data set]. Harvard Dataverse. https://doi.org/10.7910/DVN/MOVIFA
@dataset{benbrik2025arab,
author = {Ben Brik, Anis and Gilbert, Neil and Pyci{\'n}ska, Magdalena},
title = {Arab {AI} Governance Lab Dataset},
year = {2025},
publisher = {Harvard Dataverse},
doi = {10.7910/DVN/MOVIFA},
url = {https://doi.org/10.7910/DVN/MOVIFA},
note = {22 Arab League member states; CC BY 4.0}
}Tables
The tables form a star schema with F1 (22 states, 46 variables) at the centre. All tables share the ISO3 country code as key, so they merge at any level of analysis. The deposit also holds two Excel workbooks (analytical file and source archive).
| Table | Rows | Variables | Type | Key variables | Source |
|---|---|---|---|---|---|
| F1 | 22 | 46 | Cross-section | AI strategy, regulation count, maturity tier | Regulations.AI, OECD.AI, manual coding |
| F2 | 101 | 9 | Country-year panel | AI publications, Oxford score, EGDI score | OECD.AI and Scopus, Oxford Insights, UN DESA |
| F3 | 132 | 9 | WDI panel | Internet users, GDP per capita, R&D expenditure | World Bank WDI |
| F4 | 22 | 8 | Cross-section | GDP per capita 2020 to 2023, income group, fuzzy economic resources | World Bank, manual calibration |
| F5 | 137 | 7 | Oxford panel 2020 to 2025 | Overall score and three pillars | Oxford Insights (CC BY-SA 4.0) |
| F6 | 64 | 9 | EGDI biennial 2018 to 2024 | Rank, score, OSI, HCI, TII, e-participation | UN DESA |
| F7 | 22 | 6 | Cross-section | GCI, NCSI, DDL and exposure scores | ITU, e-Governance Academy |
| F8 | 176 | 21 | Merged master panel | All F-table variables | Merge of F2 to F7 |
| C1 | 22 | 10 | fsQCA fuzzy sets | Outcome and seven conditions | Calibration anchors, manual coding |
| C2 | 8 | 7 | Calibration anchors | Full-in, crossover and full-out thresholds | Ragin (2008) direct calibration |
| D1 | 21 | 4 | Qualitative profiles | Regulatory philosophy, enforcement, data sovereignty approach | Manual coding from primary documents |
| M3 | 22 | 8 | Typology matrix | Governance type, approach, enforcement, cluster label | Manual coding from D1 and F1 |
| Z1 | 26 | 7 | Codebook | Variable definitions, sources, calibration notes | Internal |
Method
With 22 cases, regression estimates are unstable and the assumptions of symmetric, additive effects are hard to defend. The Lab therefore uses fuzzy-set Qualitative Comparative Analysis (Ragin, 2008; Schneider & Wagemann, 2012), which treats governance quality as the product of combinations of conditions and allows several routes to the same outcome.
Continuous variables are converted to set membership scores (0 = fully out, 0.5 = crossover, 1 = fully in) by direct calibration against the anchors in table C2. Table C1 holds the outcome and seven calibrated conditions, among them state capacity, economic resources, digital infrastructure, data protection, international alignment and research output. Necessity and sufficiency analyses, with consistency and coverage, are deposited with the data.
The sufficiency analysis identifies three routes. Petro-fiscal surplus (UAE, Saudi Arabia, Qatar): high economic resources, state capacity and digital infrastructure. Institutional pioneer (Egypt, Tunisia): high data protection, international alignment and research output without high economic resources. Research-led emergence (Iraq): high research output without state capacity or data protection, reaching partial membership in the outcome (0.45).
Arab AI Governance Index: method
The index summarises four pillars, each on a 0 to 100 scale and weighted equally. Regulation: the Lab regulatory maturity tier, rescaled (Absent = 0, Minimal = 25, Emerging = 50, Intermediate = 75, Advanced = 100). Readiness: Oxford Insights Government AI Readiness Index, 2025. Digital government: UN E-Government Development Index 2024, multiplied by 100. Cybersecurity: ITU Global Cybersecurity Index 2024.
The score is the arithmetic mean of available pillars. A state is scored when at least three pillars are available; Comoros, Djibouti and Somalia are not scored. Mauritania, Palestine, Sudan and Yemen are scored on three pillars, which should be kept in mind when comparing them with states scored on four.
Limitations. Equal weighting is a convention, not a finding. Three pillars come from external composite indices whose own methods change between editions. Composite rankings can shift under alternative weights and normalisation choices (Ben Brik, 2026, Quality & Quantity). Pillar values are always shown beside the score so that readers can apply their own weights.
Known limitations of the 2025 release
(i) Indicator values for Comoros, Djibouti and Somalia are not displayed on this site. (ii) UAE publication counts are missing, so research rankings cover states with data only. (iii) The Oxford Insights 2025 edition revised its pillar method; changes between 2023 and 2025 combine real movement with method effects. (iv) 2024 publication counts are incomplete because of indexing lag. (v) The Cyber Exposure Index shown on the earlier site is withheld until its source is documented. (vi) Regulatory coding reflects documents identified up to the release date; legal change since then is not captured. Corrections are welcome at anis.ben.brik@usi.ch.
Sources
- ITU (2024). Global Cybersecurity Index. International Telecommunication Union.
- OECD (2024). OECD.AI Policy Observatory. oecd.ai
- Oxford Insights (2020 to 2025). Government AI Readiness Index. Annual editions. CC BY-SA 4.0.
- Ragin, C. C. (2008). Redesigning social inquiry: Fuzzy sets and beyond. University of Chicago Press.
- Schneider, C. Q., & Wagemann, C. (2012). Set-theoretic methods for the social sciences. Cambridge University Press.
- United Nations DESA (2018 to 2024). E-Government Survey. Biennial.
- World Bank (2023). World Development Indicators.
Version history
2025: first release on Harvard Dataverse. September 2026: site rebuilt with one page per theme and per state, static figures, corrected counts and a limitations section.