Bocconi UniversityProjects

Do AI Activity Indices Really Measure AI Activity?

Do AI Activity Indices Really Measure AI Activity?
  • Hugo M. V. Arsenio

    Senior Advisor

  • Riccardo Bertamini

    President · Bocconi

  • Lorenzo Caputi

    Board Chairman

  • Andrea Procopio

    Board Chairman

Abstract

AI dashboards increasingly shape economic debate, yet their headline movements often conceal what changed. For a ratio, the same rise can reflect numerator growth, denominator decline, or both. This ambiguity creates structural non-identification: adding observations or richer fixed effects cannot recover the missing information. We propose three propositions showing that a relative-index regression blends the causal effect of interest with co-movement between focal and aggregate activity and shocks shared with the outcome. These components can give the coefficient either sign under a zero effect, or erase it under a nonzero effect.

Publishing denominator levels changes what can be learned by removing the co-movement channel exactly, although shared-shock confounding remains. The problem extends to power: if most identifying variation comes from a handful of units, those units determine what the design can detect, however large the nominal sample looks. At a fixed standard error, the minimum detectable effect diverges as the effective cluster count approaches one, leaving a null result uninformative.

We apply this framework to OECD.AI’s hiring index and Eurostat employment across 24-27 European countries, where the detection floor exceeds the reported association. In a refreshed panel, that association disappears once the 2021 survey break is excluded. Ratio indices can still reveal concentration, timing, and co-movement, but establishing displacement requires more; our framework turns that boundary into a practical standard by making denominator levels a minimum disclosure requirement.

Introduction

Almost everything we know about how AI is spreading through the economy arrives as a ratio: a model family’s standing is its share of arena battles; a provider’s reach is its share of served requests; adoption is the share of surveyed firms reporting use. The most quoted labour-market indicator, the OECD.AI relative AI hiring index, is the growth of AI-skilled hiring relative to the growth of all hiring [29].

The ratio is published so that nuisances such as coverage or the survey frame cancel, and a raw numerator would be no better defined. The problem is not the ratio, but that the provider does not release the two counts. If we ask whether AI is replacing workers, or whether a model family is falling behind, we are given a number that answers neither question, and no subsequent estimator can correct it. This can be shown with the Chatbot Arena, which samples pairs of models for human comparison [54] and releases its battle logs, so both counts are observable. Across 106,134 battles among 55 models between June and August 2024 [35], the Claude family’s published share of battles fell from 34.6% to 16.4%. Over the same window its win rate against the 21 opponents present in both periods moved by 0.01-0.01 points once the family’s own model mix is held fixed, inside a bootstrap interval of ±4\pm 4 points. The primary change occurred in the sampling process: fourteen entrants accounted for 57% of late-window battles. Although daily volume remained stable, the family’s own count declined, resulting in a decreased ratio (Figure 1).

Two panels: the Claude family’s share of Chatbot Arena battles falling from 38% to 14% over the summer of 2024, beside its win rate before and after, flat at −0.01 points once opponents and own model mix are held fixed.
Figure 1. A published index and the signal it is read as. The Claude family’s share of Chatbot Arena battles halves as entrants take 57% of a flat daily volume, while its win rate is flat once the opponent set and the family’s own model mix are both held fixed [35, 27].

These examples motivate a narrow question: what can a regression on a relative activity index establish when the counts inside it are withheld? We give a formal answer, and we measure how the resulting identification and power constraints bind in the regression.

Contributions and scope

We study regressions of outcomes on published relative indices when the component levels are withheld, and we characterise the additional disclosure required for identification. Instead of estimating the employment effect of AI, we focus on whether the published indices could identify it at any size. Our contributions are:

  • We derive an exact decomposition of the coefficient from a regression of any outcome on a published relative index: the causal effect enters summed with a cyclicality term, which is the classical ratio confound, and a shock covariance term. The decomposition survives every fixed-effects design, because each within-transformation adds one equation and one unknown, so the effect is not identified from the index at any magnitude (Proposition 1).
  • We prove that publishing the denominator in levels removes the cyclicality term exactly while the covariance term survives, and that where the two have opposite signs disclosure can leave the total bias larger than before. The textbook remedy, entering the ratio alongside its parts [61], is not available to the analyst here, because the parts are proprietary; only the provider can enable it (Proposition 2).
  • We establish a lower bound on the smallest detectable effect that depends on the effective number of clusters rather than on the sample size (Proposition 3).

Empirically, we join the hiring index to Eurostat employment for 24 to 27 European countries from 2018 to 2025. A one-standard-deviation higher index is associated with about 0.5 percentage points lower employment growth, but inference varies across procedures and two or three countries carry most of the signal. In the panel reaching 2025, the association survives only through the 2021 survey break. Its minimum detectable effect is 1.8 to 4.5 times the estimate. The dashboards reveal concentration, timing and co-movement, but not displacement.

Pearson showed that indices sharing a denominator correlate even when the underlying quantities do not [60], and Kronmal restated the problem for regression: a ratio may enter only alongside its parts as main effects [61]. Kuh and Meyer, and then Belsley, developed the same point for deflated variables in econometrics [62, 63], and Aitchison’s log-ratio analysis is its compositional-data form [64]. Our published index is an Aitchison log-ratio, and Proposition 2 is Kronmal’s prescription written out for a setting where the parts are withheld.

The task framework of Autor, Levy and Murnane [13] and the robot-exposure design of Acemoglu and Restrepo [14] set the template that generative AI inherited. Experimental and firm-level work finds productivity gains on cognitive tasks [15, 16, 43]; task-based measures imply broad exposure [17, 52, 53, 24, 25, 48] against modest aggregate effects [18]; and direct employment evidence is mixed and mostly national [19, 20, 33, 42, 50, 51, 49, 44, 45], focusing on how large the effect is. We ask whether the published indicators can identify that employment effect at any magnitude. Proposition 1 shows they cannot: with only the ratio index and the outcome observed, the estimated coefficient is consistent with every value of the effect: the coefficient can take either sign when the true effect is zero, and be zero when it is not.

Our object is the measurement layer these dashboards constitute [56]: OECD.AI’s LinkedIn-derived hiring, migration and skills series [29], venture-capital aggregates [30], enterprise adoption surveys [32] and vacancy rates [47]. The same form governs how AI systems themselves are watched: Chatbot Arena publishes shares of pairwise battles [54, 35], and Singh et al. [34] document unequal sampling, selective disclosure of private variants and silent deprecation behind its ratings. It establishes that one platform’s numbers are distorted, whereas we show that a published ratio cannot carry a causal reading however well the platform behaves.

When many decision-makers run the same algorithm, its mistakes stop cancelling out and they share the same errors [57]. The same logic applies to measurement: one dashboard supplies most public claims about AI and jobs, so having many readers gives no independent check on it. Separately, a prediction that gets acted on changes what it predicted [58], and people who know they are being scored move to improve the score [59]. A ratio is especially easy to influence, since it rises either when AI-skilled hiring goes up or when total hiring goes down.

Cluster-robust asymptotics fail when a handful of clusters carry the identifying variation [21, 22, 38], effective cluster counts diagnose it [55], and randomization inference trades the reference distribution for an exchangeability assumption [36, 37]. The shift-share literature supplies the analogous warning for exposure designs [41, 40], as does pre-testing for parallel trends [39].

Methods

Data

Three providers supply the labour-market analysis, OECD.AI, Eurostat and the ILO, and Table 1 gives the series and their coverage, including a policy score of our own construction. The arena exhibit uses a further source, the released Chatbot Arena battle logs [35], which the table omits. We exclude 2026 everywhere as a partial year.

Table 1: Primary sources and their coverage. The text below states what each is used for.

Source and variable

Coverage

OECD.AI relative AI hiring index [5]

2018-01 to 2026-01, monthly, averaged to years

OECD.AI AI skills migration [3]

Per 10,000 LinkedIn members; extracts spanning 2019-2024 (Panel A) and 2021-2025 (Panel B)

OECD.AI VC in AI start-ups [1, 2]

By country and by industry, 2012-2026; the industry table’s country dimension is the investor’s country

OECD.AI AI skills penetration [4]

By country and industry, 2016-2024

Eurostat employment by occupation [6, 65]

Ages 15-64; Panel A uses an extract of table lfsa_egai2d, the other samples lfsa_epgais over 2016-2025

ILO generative-AI exposure scores [48]

427 ISCO-08 unit groups, aggregated to major groups with Eurostat employment weights

Eurostat employment by NACE two-digit [7]

Ages 20-64, 2018-2025

Eurostat covariates [8, 9, 10, 11, 12] and own policy score

Country-year; policy score hand-coded 0-3 (one point each: national AI strategy in force, major update, EU AI Act in force)

No single join spans 2019 to 2025, so we build two samples. Panel A (24 countries, 2019 to 2024, 144 country-years) has the longer window and the full covariate set and is the primary test; Panel B (26 countries, 2021 to 2025, 130 country-years) is the only sample that observes 2025 and the only one industry investment joins to. They overlap in 88 of Panel B’s 130 country-years but differ in window, covariate set and index extract, so Panel B tests Panel A’s coefficient on a different extract rather than resampling it, and a pooled panel would have concealed the disagreement. The occupation designs use no migration series, which is why they reach 27 countries.

Identification

Dashboards report activity as a ratio. Let Ai,tA_{i,t} be a focal count and Bi,tB_{i,t} an aggregate that contains it, both observed on one platform, and let the published series be

Ri,t=ΔlogAi,tΔlogBi,t.R_{i,t} = \Delta \log A_{i,t} - \Delta \log B_{i,t}.

For the hiring index AA is AI-skilled hires and BB total hires; for a leaderboard AA is one family’s battles and BB all battles. This object has two key properties.

  • Invariant to whatever scales AA and BB together, which is exactly why providers publish it: platform coverage, the member base and any common reporting convention cancel.
  • Level information is discarded by the invariance: the same value may reflect a collapsing numerator, a growing denominator, or both, and the series cannot distinguish among them.

Two departures from this idealised object matter.

  • The published index is smoothed, so R=ΔlogAΔlogB+ωR = \Delta \log A - \Delta \log B + \omega, where the slack ω\omega is one further free nuisance: it leaves Proposition 1 untouched but weakens Proposition 2, since disclosing BB then recovers the AI-specific shock only up to ω\omega.
  • This instrument is published as a growth differential, whilst most others of this kind are published as level shares S=A/BS = A/B. The two are one object, since ΔlogS=ΔlogAΔlogB\Delta \log S = \Delta \log A - \Delta \log B, so a regression on ΔlogS\Delta \log S inherits the decomposition term for term, and one on ΔS\Delta S differs by a unit- and period-specific rescaling that changes no term’s status.

Consequently, the results apply whether the provider publishes a level share or a growth differential.

For the hiring index we can argue that the two counts move for unrelated reasons but cannot show it, because neither is published. By contrast, the same analysis is possible for Chatbot Arena because it releases its battle logs [54]. The figures reported in the introduction locate the change in the numerator. Daily battle volume remained flat, and the roster fell from 41 models to 38, ruling out an expansion of BB. Yet observing BB would establish only that the movement occurred through AA, not why AA moved, as Proposition 2 shows. And holding opponents fixed, but not the family’s own model mix, yields +2.5+2.5 points (95% CI 0.4-0.4 to +5.5+5.5) rather than 0.01-0.01, showing that the performance measure is sensitive to changes in within-family composition (Figure 1).

The leaderboard’s headline is a Bradley-Terry rating, and it does not escape the problem, since it conditions on a pairing policy and prompt distribution the platform does not publish [34]. The required disclosure is the analogous one: per-family battle counts, the eligible model set per period, and the sampling policy, in levels.

We develop the algebra for the hiring case because the employment data needed to evaluate displacement are available and the index’s denominator enters directly into employment dynamics. The identification problem therefore arises from the index itself rather than from a lack of outcome data. The identity Ei,t=Ei,t1+Hi,tSi,tE_{i,t} = E_{i,t-1} + H_{i,t} - S_{i,t} makes aggregate hiring a common component of both the index and employment growth gi,tg_{i,t}, while the numerator measures the AI-skilled hiring margin whose effect on employment is the parameter of interest. Assume finite second moments and a common relative-cyclicality coefficient across the units entering the pooled moments below. Write the within-country projection of AI-skilled on total hiring as Δlogh=γΔlogH+u\Delta \log h = \gamma\, \Delta \log H + u, where γ\gamma is a relative cyclicality parameter and uu an AI-specific shock orthogonal to ΔlogH\Delta \log H by construction, and let

gi,t=δui,t+ηi,t,g_{i,t} = \delta\, u_{i,t} + \eta_{i,t},

where δ\delta is the displacement parameter of interest and η\eta collects everything else, including the cycle. We define δ\delta outside this equation, as the causal effect of intervening on the AI-specific hiring shock while holding η\eta, and with it aggregate hiring, fixed; the equation is a linearity restriction on that structure, not a definition of η\eta, which would make it an accounting identity holding for any δ\delta. We place no restriction on Cov(u,η)\operatorname{Cov}(u, \eta): a common shock can move AI-specific hiring and employment together without either causing the other. Substituting the projection into the index gives R=(γ1)ΔlogH+uR = (\gamma - 1)\Delta \log H + u, and hence

Cov(R,g)Var(R)β, estimated  =  (γ1)Cov(ΔlogH,g)Var(R)relative cyclicality  +  δVar(u)Var(R)displacement  +  Cov(u,η)Var(R)common shock\underbrace{\frac{\operatorname{Cov}(R, g)}{\operatorname{Var}(R)}}_{\beta,\ \text{estimated}} \;=\; \underbrace{\frac{(\gamma - 1)\operatorname{Cov}(\Delta \log H,\, g)}{\operatorname{Var}(R)}}_{\text{relative cyclicality}} \;+\; \underbrace{\frac{\delta \operatorname{Var}(u)}{\operatorname{Var}(R)}}_{\text{displacement}} \;+\; \underbrace{\frac{\operatorname{Cov}(u, \eta)}{\operatorname{Var}(R)}}_{\text{common shock}}

Our only object of interest is the displacement term.

Proposition 1 (Non-identification). Let β=Cov(R,g)/Var(R)\beta = \operatorname{Cov}(R, g)/\operatorname{Var}(R), and let the data identify only β\beta and Var(R)\operatorname{Var}(R). Then δ\delta is not identified from (R,g)(R, g):

δR, bR,γR, c=Cov(u,η)R:the decomposition holds with β=b.\begin{gathered} \forall\, \delta \in \mathbb{R},\ \forall\, b \in \mathbb{R}, \quad \exists\, \gamma \in \mathbb{R},\ \exists\, c = \operatorname{Cov}(u, \eta) \in \mathbb{R} : \\[2pt] \text{the decomposition holds with } \beta = b. \end{gathered}

In particular β\beta takes either sign with δ=0\delta = 0, and β=0\beta = 0 is attainable with δ0\delta \neq 0. The statement is unchanged under any linear within-transformation MM, including country and year fixed effects and country-specific trends.

Proof. Write CHg=Cov(ΔlogH,g)C_{Hg} = \operatorname{Cov}(\Delta \log H, g) and V=Var(R)V = \operatorname{Var}(R), so the decomposition reads βV=(γ1)CHg+δVar(u)+c\beta V = (\gamma - 1)C_{Hg} + \delta \operatorname{Var}(u) + c: one equation, three unknowns (γ,δ,c)(\gamma, \delta, c). Given any target (δ,b)(\delta, b), set γ=1\gamma = 1 and c=bVδVar(u)c = bV - \delta \operatorname{Var}(u); then β=b\beta = b. Taking δ=0\delta = 0, c=bVc = bV gives β=b\beta = b of either sign; taking b=0b = 0, c=δVar(u)0c = -\delta \operatorname{Var}(u) \neq 0 gives β=0\beta = 0 with δ0\delta \neq 0. Equivalently, any orthogonal split of RR is admissible, and δ\delta rescales freely as Var(u)0\operatorname{Var}(u) \to 0. For the fixed-effects claim let MM be symmetric and idempotent, removing country and year means or any further within-country regressors. Then MR=(γ1)MΔlogH+MuMR = (\gamma - 1)M\Delta \log H + Mu and Mg=δMu+MηMg = \delta Mu + M\eta, so

βM=(γ1)Cov(MΔlogH,Mg)+δVar(Mu)+Cov(Mu,Mη)Var(MR),\beta_M = \frac{(\gamma - 1)\operatorname{Cov}(M\Delta \log H,\, Mg) + \delta \operatorname{Var}(Mu) + \operatorname{Cov}(Mu,\, M\eta)}{\operatorname{Var}(MR)},

which is the same decomposition with transformed moments and a new free nuisance Cov(Mu,Mη)\operatorname{Cov}(Mu, M\eta). Each MM buys one equation and one unknown, and adds no restriction unless Cov(u,η)\operatorname{Cov}(u, \eta) is itself restricted.

Three restrictions bear on the proposition; the first two jointly overturn it, the third alone suffices:

  • γ=1\gamma = 1, so AI-skilled and total hiring share a cycle exactly, eliminates the first term, but it is an assumption about the very quantity the index conceals and no published series tests it.
  • Cov(u,η)=0\operatorname{Cov}(u, \eta) = 0, asserting that nothing moving employment also moves AI-specific hiring; however, European hiring cooled and AI-specific hiring fell together over 2022 and 2023.
  • Using an instrument that shifts uu while moving neither η\eta nor the withheld denominator ΔlogH\Delta \log H. We find none in the published data; the likeliest candidates, venture capital and skills migration, may affect employment through channels other than the AI-specific hiring shock, and both are procyclical, so they cannot be assumed to shift uu alone.

The related-work section places this decomposition in the ratio literature, leaving the sign of the cyclicality term open. The sign of γ\gamma is not determined a priori: adjustment costs make specialist hiring lumpy, arguing for γ<1\gamma < 1 and a negative β\beta with no displacement at all, while specialist hiring fell further than aggregate hiring in 2022 and 2023, arguing for γ>1\gamma > 1. Published data cannot separate the two. Figure 2 shows how, with δ\delta held at zero, the coefficient changes sign once Cov(u,η)\operatorname{Cov}(u, \eta) crosses (γ1)Cov(ΔlogH,η)-(\gamma - 1)\operatorname{Cov}(\Delta \log H, \eta). Panel (a) is not calibrated to the data, so it establishes only that the bias can run in either direction.

Proposition 2 (What disclosure buys, and what it does not). Suppose the provider also publishes ΔlogB\Delta \log B, here total platform hiring. Then γ\gamma and uu are identified, and the coefficient on uu in the projection of gg on (ΔlogH,u)(\Delta \log H, u) equals δ+Cov(u,η)/Var(u)\delta + \operatorname{Cov}(u, \eta)/\operatorname{Var}(u). The relative-cyclicality channel is removed exactly and γ\gamma becomes measurable; δ\delta remains confounded by Cov(u,η)\operatorname{Cov}(u, \eta) and is still not identified.

Proof. Given RR and ΔlogH\Delta \log H, the series Δlogh=R+ΔlogH\Delta \log h = R + \Delta \log H is observed, so γ\gamma is the projection coefficient of one observed series on another and uu is its residual. Regressing gg on (ΔlogH,u)(\Delta \log H, u) with uΔlogHu \perp \Delta \log H gives a coefficient on uu of Cov(u,g)/Var(u)\operatorname{Cov}(u, g)/\operatorname{Var}(u), which equals δ+Cov(u,η)/Var(u)\delta + \operatorname{Cov}(u, \eta)/\operatorname{Var}(u).

Two panels: the published-index coefficient crossing zero as the covariance between the AI-specific shock and the outcome varies, and the minimum detectable effect plotted against the effective cluster count with the reported estimate of 0.47 marked.
Figure 2. (a) With δ = 0 everywhere, the published-index coefficient changes sign as Cov(u, η) varies; disclosure removes the cyclicality channel only, and over part of the range sits further from the truth. (b) The Proposition 3 detection floor at σ = 0.2405: the minimum detectable effect at 80% power against the effective cluster count, with this design marked at G* = 2.3 (floor 2.13) and the reported estimate 0.47 dashed.

The removal is exact because uu becomes the residual of an observed projection. Disclosure does not however shrink the bias: where the channels have opposite signs, removing one enlarges the net error, as Figure 2 shows, so a disclosed index can sit further from the truth than the ratio it replaced while being far more informative about why. One caveat is substantive: the provider publishes rates, not counts, so recovering ΔlogH\Delta \log H also needs the member-base growth that cancels out in the ratio. The required disclosure is therefore the underlying levels. Even with those levels, however, identification is not the only constraint.

Proposition 3 (Detection floor). For a clustered design with standard error σ\sigma and GG^{*} effective clusters, the smallest effect detectable at the 5% level with power 1κ1 - \kappa is approximately

MDE=(t0.975,G1+t1κ,G1)σ,\mathrm{MDE} = \left(t_{0.975,\, G^{*}-1} + t_{1-\kappa,\, G^{*}-1}\right) \sigma,

which diverges as G1G^{*} \to 1 at fixed σ\sigma, irrespective of the number of observations.

GG^{*} is the effective cluster count implied by the concentration of the regression’s own influence: it falls below the nominal count whenever a few clusters carry most of the squared score mass, so it is a property of the design: adding observations to the same few clusters does not raise it; only variation spread across more clusters would, and the panel already spans the continent.

Results

Dashboards

Financing is deal-concentrated, though less than the aggregate share implies. IT infrastructure takes 65.21% of European AI venture capital in 2025, but one investor-country cell carries most of that, and the share falls to 27.8% on EU27 member rows, where there is no rebound above the 2021 level. One robust finding remains: mega-deals were about 73% of global AI investment value [30].

The published index increases in 2024, eighteen months after the ChatGPT release [28]. We hold two extracts of the index: the medians here use the refreshed one and Panel A’s regressions the other, so agreement between them is corroboration only within one data ecosystem. The median across 27 countries ran 3.30, 1.48, 3.43 over 2018 to 2020, turned negative at 1.31-1.31, 4.93-4.93, 4.32-4.32 over 2021 to 2023, then rose to 6.80 and 7.96. The upturn is common to both extracts and survives changes of country set, means for medians, and no year-averaging. The pre-2021 level does not, so the rise partly recovers the earlier level rather than establishing a new one.

Capital and talent co-move, and employment reallocates. Lagged venture capital goes with a higher hiring index the following year under country and year effects (3.60, SE 1.73, p=0.048p = 0.048), but not without them and not on the industry panel, and the outcome in that model is the index itself, so Proposition 1 bounds it too. Employment shares shift toward IT infrastructure and healthcare and away from the residual. These are shares, so they describe reallocation, and the programming trend predates generative AI [31].

Treating November 2022 as the event date, three of four annual designs reject their pre-trend test; the fourth, whose outcome is employment growth rather than the index, does not (p=0.100p = 0.100) and shows nothing after the date. A shift-share sensitivity agrees, since its bivariate first stage flips sign once controls enter [41, 40, 39].

AI hiring and employment growth

Proposition 1 says a coefficient of either sign is compatible with no displacement. We nonetheless estimate the regression and assess what it can support. The outcome Yi,tY_{i,t} is total employment growth in percent year on year, and the baseline is

Yi,t=αi+λt+β1Hiringi,t+β2Migrationi,t+β3Tertiaryi,t+β4Policyi,t+εi,t,Y_{i,t} = \alpha_i + \lambda_t + \beta_1 \mathrm{Hiring}_{i,t} + \beta_2 \mathrm{Migration}_{i,t} + \beta_3 \mathrm{Tertiary}_{i,t} + \beta_4 \mathrm{Policy}_{i,t} + \varepsilon_{i,t},

with errors clustered by country. Twenty-four clusters is too few for reliable asymptotics, so we fix four procedures in advance: CRV1 intervals (t(23)t(23)) [see the note below]; a wild cluster bootstrap-t with Rademacher weights, null imposed, B=99,999B = 99{,}999, and seed 42 [21, 22]; the same tt referred to the effective cluster count; and studentized randomization inference [36, 37]. A 200-variant specification curve leaves 198 estimates negative, but 192 of the 200 perturb one 144-observation regression.

Table 2: Robustness of the relative-AI-hiring coefficient (per 1 SD), Panel A. Intervals are CRV1 t(23); p columns are analytic clustered and wild cluster bootstrap-t (99,999 draws for the baseline, 4,999 elsewhere, one shared draw sequence).

Check

Coefficient

95% interval

p (analytic)

p (bootstrap)

Two-way FE, baseline

−0.47

[−0.97, +0.02]

0.061

0.037

+ GDP-growth and unemployment controls

−0.46

[−0.83, −0.09]

0.017

0.014

Without COVID years (n = 96)

−0.45

[−0.99, +0.09]

0.097

0.003

Without 2021, the LFS break year (n = 120)

−0.45

[−1.07, +0.18]

0.157

0.030

Relative AI hiring is the only baseline regressor with a recurring partial association: a 1 SD higher index goes with roughly 0.3 to 0.6 percentage points lower employment growth within a country, while migration, tertiary share and policy score show none. Table 2 reports four specifications; the coefficient stays negative throughout, but the 95% intervals reach or cross zero in most rows. Employment weighting strengthens it to 0.89-0.89 per SD (p=0.004p = 0.004) and country trends shrink it to 0.34-0.34 (p=0.17p = 0.17). Adding GDP growth and unemployment leaves it unchanged: the channel Proposition 1 formalises runs through total hiring, which neither control measures.

The identifying variation is concentrated in very few clusters based on the decomposition: Türkiye carries 47% of the squared mass on the restricted scores the bootstrap resamples, giving 3.2 effective clusters, while on the unrestricted scores Estonia carries 61% and G=2.3G^{*} = 2.3. Estonia and Latvia hold the panel’s two largest regressor values, +4.6+4.6 and +4.4+4.4 standard deviations in 2019, so small-country volatility affects hiring as well as migration. Leave-one-cluster-out agrees: dropping Türkiye moves the baseline from 0.47-0.47 (p=0.061p = 0.061) to 0.35-0.35, Latvia to 0.50-0.50 and Estonia to 0.73-0.73 (p=0.005p = 0.005), so single-cluster deletions span 0.35-0.35 to 0.73-0.73.

As the estimate depends on the cluster, the p-value depends on the procedure. Each bootstrap p-value in Table 2 undercuts its analytic counterpart, in one row by 0.127, while referring the same tt to t(G1)t(G^{*} - 1) gives p=0.25p = 0.25 unrestricted and p=0.17p = 0.17 restricted. It might seem that the bootstrap over-rejects and t(G1)t(G^{*} - 1) corrects it; however, simulated under the null on this design’s own matrix, both the bootstrap and CRV1 are correctly sized, so their gap in Table 2 is unexplained, while t(G1)t(G^{*} - 1) rejects in none of the 353 replications as concentrated as the data, so p=0.25p = 0.25 is uninformative, which restates Proposition 3. Two cautions about GG^{*}: it is estimator-dependent (5.5 under an alternative weighting, p=0.11p = 0.11) and noisy (median 3.4, 5 to 95% range 1.8 to 7.2). Studentized randomization inference [36, 37] gives p=0.045p = 0.045 over 20,000 draws and is the least sensitive of the four to the error process, though the exchangeability it rests on strains this panel, since the heterogeneity driving the leave-one-out range is what a permutation null denies. Across procedures, the baseline runs from 0.037 to 0.25, but the 0.25 end is uninformative: it comes from a reference distribution on which this design could never have rejected. We do not need the estimate to be insignificant: Proposition 1 bars the displacement reading at every value.

At the reported σ=0.2405\sigma = 0.2405 and GG^{*} between 2.29 and 5.5, the Proposition 3 floor at 80% power runs from 2.13 down to 0.86, and to 0.70 with all 24 clusters effective (Figure 2b); every value exceeds the 0.47 the primary specification reports. Inverting the noncentral tt gives exactly 2.38 and 0.87, so the closed form understates the constraint at the concentrated end. The comparison is not fixed in advance, since σ\sigma and GG^{*} are both estimated from the regression being evaluated, so a floor above the estimate is implied by p>0.05p > 0.05. That ratio, 1.8 to 4.5 times the estimate, therefore rescales the null result rather than testing it. Against a benchmark fixed in advance, even the floor’s best case, 0.70 points of annual employment growth per standard deviation, exceeds any employment effect documented for earlier automation waves [14], so effects of every previously observed size were undetectable by design.

The natural probe of Proposition 1 is a measure whose denominator is not total hiring [32, 46, 47]. AI talent concentration, whose denominator is platform membership, is published in three encodings on the same 221 observations. Only one has power against the baseline effect (90%), and it is null: +0.02+0.02 per SD, p=0.91p = 0.91. The other two are underpowered at 70% and 13%, so their negative point estimates are not evidence, and two further probes cannot discriminate at this sample size either. The second check is Panel B: with predictors in their own units and venture capital as log(1+deals)\log(1 + \text{deals}), the hiring association holds in sign but weakens to 0.027-0.027 under two-way FE (p=0.059p = 0.059, bootstrap 0.056), about 0.24-0.24 per SD in Panel A. Every 2021 observation carries Eurostat’s break flag from the Labour Force Survey redesign, and 2021 growth is measured against the 2020 depression; excluding 2021 moves the coefficient to +0.002+0.002 (t=0.07t = 0.07) and excluding all break-flags to +0.003+0.003, while the same exclusion leaves Panel A at 0.45-0.45.

The two extracts of the index are not the same series. On the 208 country-years where they overlap they correlate 0.447; one is positive in 200 of those cells and the other in 119, so they disagree in sign on at least 81. This undocumented divergence is itself evidence: under the two-way demeaning the baseline uses, the extracts correlate 0.528, so the identifying variation behind the 0.47-0.47 is only half-reproducible.

Local projections [23] give cumulative associations between 0.36-0.36 and 0.62-0.62 over horizons 0 to 3, with only h=2h = 2 excluding zero under CRV1, and the lead estimate is imprecise but larger than the contemporaneous one, so reverse causality is not excluded.

Occupations

A displacement hypothesis predicts losses concentrated in occupations more exposed to AI, which the country regressions cannot see. We test it on a country × occupation × year panel of nine ISCO-08 major groups, regressing occupation employment growth on the hiring index interacted with exposure, under country × year, country × occupation and occupation × year fixed effects, clustered by country. Exposure is the index of the published ILO Working Paper [48], 427 unit groups aggregated to nine with Eurostat employment weights; an unweighted crosswalk and three author-coded rank proxies [24, 25] are robustness encodings.

The country × year effects absorb every country-level shock in levels, so identification is cross-occupation within country-year, but they do not absorb the identification channel above: substituting the decomposition gives Ri,tXo=(γ1)ΔlogHi,tXo+ui,tXoR_{i,t} X_o = (\gamma - 1)\Delta \log H_{i,t} X_o + u_{i,t} X_o, and a country-year term times an occupation characteristic survives those fixed effects exactly as the displacement term does. Exposure indices load on cognitive and clerical content, where cyclical sensitivity also differs.

The interaction is negative across five exposure encodings and an employment-weighted variant, ranging from 0.34-0.34 to 1.69-1.69. Two pass the 5% threshold under CRV1 and one under the wild bootstrap. Excluding the COVID years, the largest economy, or 2021 preserves the sign. With G=4.7G^{*} = 4.7, referring the statistic to t(G1)t(G^{*} - 1) gives p=0.16p = 0.16.

Fourteen of the 1,458 observations show annual growth above 20 log points, mostly 2021 survey-redesign values; trimming them halves the published-index estimate. The occupational levels, however, run in the opposite direction: growth rises monotonically with exposure in both eras, with the era-over-era gain largest for the least exposed tercile, so the negative interaction is a within-country-year co-movement with the index.

Limitations

Relative AI hiring measures how fast AI-skilled hiring changes relative to overall hiring, not the number of AI jobs. Skills migration is per 10,000 LinkedIn members, so small-country rates are volatile. The venture-capital data measure financing activity, so large deals can dominate.

Proposition 1 formalises one mechanism under a linear projection. It does not model platform selection or measurement error, and the published index is a smoothed percentage change, so Proposition 2 describes an idealised series. Our δ\delta is a direct effect holding η\eta, and so aggregate hiring, fixed: displacement operating through total hiring loads onto γ\gamma, so Proposition 2 is not a route to a total effect, which is not identified either.

Net employment cannot reveal gross displacement [26]: separations, hours, wages and task reallocation may move while headcount does not. Every AI-activity measure here derives from LinkedIn’s member base, and the arena exhibit from a second platform with its own sampler, so platform composition can move the indicators independently of the quantity of interest. Our nulls are failures to reject.

Discussion

This work asked what a published activity ratio can identify. A regression of any outcome on the index returns the sum of three terms, only one of which is the causal effect, and no fixed-effects design separates them (Proposition 1). Publishing the denominator in levels removes the cyclicality term exactly and leaves the shared shock in place (Proposition 2), and the effective number of clusters bounds what any such design could detect (Proposition 3). Both exhibits show these limits binding: the Arena share halved with the win rate flat because the sampling changed, and the employment association rests on two or three countries, sits below its own detection floor, and reproduces on the 2025 panel only through the 2021 survey break. The dashboards do support concentration, timing and co-movement.

More broadly, the decomposition applies wherever an ecosystem is watched through ratios, and each result names the disclosure that would answer it. Proposition 1 sets the disclosure minimum: both counts in levels, at the frequency of the ratio, including the base whose growth cancels in the published rate. The Arena exhibit adds the sampling layer: which units were eligible in each period and how they were sampled, without which a share confounds performance with exposure. The results add versioning: the divergence we document arrived with no published notice, so extracts should be citable by version. Proposition 3 then obliges the analyst rather than the provider: report the effective number of clusters and the smallest detectable effect alongside any estimate built on such an index.

This post carries the paper’s main body. Appendices A to D — dashboard descriptives and occupation results, auxiliary tests and inference diagnostics, the specification curve, and the source and sample tables — are in the attached PDF.

Resources

1 file