Methods · Net Zero Industrial Policy Lab

Narrowing the green dictionary

An association screen for technology-specific HS-6 trade lines, and why it did not work.

Discovery power

0.98

28 non-anchor codes clear the gate against 28.5 expected by chance. The screen adds no detectable signal.

False positives

3.4–13.6%

Share of the 4,768 codes in no chain at all that the decision gate accepts.

Narrowing that held

40

Core codes across eight chains, from $4,279B asserted. Produced by reading HS headings, not by the screen.

Chains unmeasurable

3

EVs, geothermal and electrolyzers have no HS-6 line that names them.

An association screen for technology-specific HS-6 lines

Draft, 2026-08-14. Gilberto García-Vazquez, GripPoint LLC, for the Net Zero Industrial Policy Lab. Code and reproduction in the sibling methodology/ directory; every table regenerates from raw inputs via methodology/reproduce.sh.

§1

Summary

The NZIPL green dictionary maps HS-6 trade lines to eleven clean technology chains. It asserts membership without any weight that separates a line specific to a technology from a generic industrial line swept in alongside it. The consequence is measurable. Wind's basket is 89 codes and $1,036B of 2024 world trade, of which the only code whose heading names wind is 0.64%. Its largest lines are copper ore, miscellaneous plastics and motor-vehicle gearboxes. Ranked on that basket, Chile is the world's 7th largest wind exporter and Peru the 9th.

This document set out to specify a statistical screen for that problem, built it, tested it, and reports a negative result on the screen alongside a positive result on what actually worked.

The narrowing works. The screen does not. Adjudicating 353 codes against the legal text of their HS headings cut eight chains from $4,279B of asserted trade to 40 core codes, and every surviving leaderboard reads as the industry: nuclear becomes Kazakhstan, Russia, Canada, France and Namibia; wind becomes China, Germany, Denmark, India and Spain; magnets becomes China 61%, Japan, Vietnam and the Philippines. The asserted baskets had ranked Chile 1st in magnets, Australia 1st in nuclear and 2nd in batteries, all on ore.

The screen contributed nothing measurable to that. Section 8.5 builds the null the first draft lacked: score the 4,768 HS-6 lines the dictionary places in no chain against each chain's own roster. The decision gate fires on 3.4% to 13.6% of them. And net of the anchor codes the method was handed, 28 non-anchor basket codes clear the gate against 28.5 expected by chance, a lift of 0.98. Only nuclear (3.22) and biofuel (1.76) beat chance; solar (0.19), magnets (0.44), batteries (0.80) and wind (0.92) are at or below it, and heat pumps produces none at all.

Three further defects are structural rather than tunable. The anchors are drawn from the dictionary under test and define the RCA vector they are then scored against, so their high scores are mechanical; leave-one-out destroys them where it can be computed at all (biofuel's biodiesel anchor falls 6.10 to 0.66). The stated null of 1.0 is a definition the data do not honour, with median null M2 running 0.60 in nuclear to 1.22 in wind. And M1's attainable ceiling varies eighteen-fold across chains, so one threshold is a different test in each.

What survives is worth keeping. The adjudication procedure in Section 5.5, which reads a heading's legal scope and asks whether it can lawfully contain the technology, produced every result in Section 7 and is reproducible by anyone with the nomenclature. Two validity gates in Section 8.4 diagnose when a roster has degenerated into a proxy for one country. And the exercise surfaced five defects no weighting scheme addresses, including a vintage collision that double-counts $7.32B into wind and solar by construction, a transmission basket missing roughly $40B of the goods that define it, and an electrolyzer basket that is 34% semiconductor manufacturing equipment.

The honest recommendation is in Section 12: adjudicate, do not screen.

§2

The problem, measured

2.1 What the dictionary asserts

green_dictionary.csv carries 715 rows across eleven technologies. Each row maps one HS code to one technology with a role (Raw Material, Processed Material, Product Component, Process Equipment, Final Product), a stage, and a split_weight.

split_weight splits within a technology, across duplicate role or position rows. It never splits across technologies. Summing it per code over all its rows returns 1.0 for 324 codes, 2.0 for 71, 3.0 for 18, and up to 10.0. A code claimed by ten chains is claimed at full value by each of them.

2.2 Three distinct symptoms

Dual-use components. Wind's basket contains HS 8708.30 (motor-vehicle brakes, $37.6B) and HS 8708.40 (motor-vehicle gearboxes, $69.4B), the latter tagged downstream, which asserts that car gearboxes are near-final wind product. HS heading 8708 is legally "parts and accessories of the motor vehicles of headings 8701 to 8705". A wind turbine is not a motor vehicle, so the wind content of those $107B is zero by construction of the nomenclature.

Raw material over-attribution. Copper ore ($104.5B) is claimed by nine chains at full value. Nickel ore by ten. Iron ore by five. Alumina by seven.

Cross-chain double count. The eleven baskets sum to $6,946B against $3,772B of unique trade, an inflation of 84%.

T1. Multi-chain membership and the inflation it produces

chains claiming a codecodesworld 2024 trade
1307$2,549B
277$586B
326$236B
45$16B
58$246B
61$0B
71$18B
92$116B
101$5B

Unique trade the dictionary touches: $3,772B. Sum of the eleven baskets: $6,946B. Inflation 84%.

T2. Codes claimed by four or more chains

hs6world 2024chainsdescription
260111$150.9B5Iron ores and concentrates: non-agglomerated
260300$104.5B9Copper ores and concentrates
847989$54.2B5Machines and mechanical appliances: having i
281820$18.0B7Aluminium oxide: other than artificial corun
260112$17.0B5Iron ores and concentrates: agglomerated (ex
260600$11.8B9Aluminium ores and concentrates
841989$11.1B5Machinery, plant and laboratory equipment: f
847982$6.8B4Machines: for mixing, kneading, crushing, gr
260200$6.5B5Manganese ores and concentrates, including f
260400$5.0B10Nickel ores and concentrates
261400$3.8B4Titanium ores and concentrates
261390$3.7B5Molybdenum ores and concentrates: other than

T3. The wind basket today, top 12 lines by world trade

89 codes, $1,036B of 2024 world trade.

hs6positionworld 2024% of basketdescription
260300upstream$104.48B10.1%Copper ores and concentrates
392690midstream$81.71B7.9%Plastics: other articles n.e.c. in chapter
870840downstream$69.35B6.7%Vehicle parts: gear boxes and parts thereo
732690midstream$57.75B5.6%Iron or steel: articles n.e.c. in heading
854370midstream$54.58B5.3%Electrical machines and apparatus: having
847989midstream$54.18B5.2%Machines and mechanical appliances: having
760110midstream$37.69B3.6%Aluminium: unwrought, (not alloyed)
870830midstream$37.64B3.6%Vehicles: parts, mounted brake linings
390210midstream$26.16B2.5%Propylene, other olefin polymers: polyprop
390110midstream$25.84B2.5%Ethylene polymers: in primary forms, polye
850300downstream$24.72B2.4%Electric motors and generators: parts suit
853650midstream$22.13B2.1%Electrical apparatus: switches n.e.c. in h

The only code whose heading names wind, 850231, is 0.64% of the basket ($6.64B).

2.3 Why it matters

The station's headline, its chain rankings and the Atlas's predicted-competitiveness scores all read off these baskets. Under the current wind basket, world wind exports are $1,036B and the top ten includes Chile, Peru, Mexico and Korea. Under a screened core the total is $25.3B and the top ten is China, Germany, Denmark, India, Spain, the United States, Turkey, Italy, France and Portugal.

T4. World wind exports 2024: current basket vs a screened basket

rankcurrent basket$Bsharescreened core$Bshare
1CHN165.416.0%CHN6.5025.7%
2DEU104.710.1%DEU3.3513.2%
3USA84.68.2%DNK2.218.7%
4JPN48.64.7%IND2.148.4%
5MEX35.73.4%ESP1.767.0%
6ITA35.33.4%USA1.375.4%
7CHL33.13.2%TUR1.134.5%
8KOR30.02.9%ITA0.562.2%
9PER24.92.4%FRA0.522.1%
10FRA24.12.3%PRT0.491.9%

Basket total: $1,036B current, $25.3B screened. The current basket ranks Chile 7th and Peru 9th in world wind exports, on copper ore.

§3

What the task requires

Three questions are routinely conflated into the single share column, which today holds 1.0 for 634 of 646 projected rows.

  1. Membership. Is this code in chain c at all?
  2. Attribution. What fraction of world trade in this code is clean-relevant, as against automotive, construction and general industrial use?
  3. Allocation. Of the clean part, how does it divide across chains that claim it?

A bill of materials speaks to (1). This screen addresses (1) and bounds (2). Section 7.3 reports the finding that (3) largely dissolves once (1) is resolved, because the codes that force an allocation rule are the codes that survive no chain's screen.

§4

Data

inputvintagerole
green_dictionary.csvCVCE, pinned blob verified against ecoclassical/NZIPL-CVCE HEADmembership claims under test
BACI HS12 V2026012024, 11,109,411 recordstrade values and exporter identity
products.csvprojector outputdictionary codes walked into HS12 trade space
country_hs6_all_2012_2024_v1.csvderivedper (economy, HS-6) exports
out/baci_2024_totals.csvderived, this worktotal merchandise exports per economy, 226 economies, world $22,885B

The last file is new and load-bearing. Balassa RCA requires total merchandise trade in the denominator. Earlier drafts used the 428-code clean basket, which asks whether an economy is specialised among clean goods and so presupposes the classification under test. Section 6, Alternative 6 documents the difference: India's wind RCA falls from 4.2 to 3.1 when the denominator is corrected, which removes it from the roster.

Nothing is fetched at runtime. All inputs are on disk and gitignored.

§5

Method

5.1 Anchors

An anchor is an HS-6 line a customs officer can only be classifying if the good really is that technology. The anchor set defines what the chain is, against which every candidate is measured.

Candidates are drawn from dictionary rows with role == 'Final Product' and then tested individually, because that label is necessary and not sufficient. Geothermal's Final Product rows are steam turbines, boring machinery and steel well casing, which are oil-and-gas codes. A polluted anchor poisons the roster and therefore every score downstream of it, so anchor validation is the step with the highest leverage in the method and the one least amenable to automation.

Anchor quality is reported per chain as clean, weak, or none. none is a legitimate result.

5.2 Rosters

For chain c with anchor basket A, Balassa RCA over total merchandise exports:

> RCA(i,c) = ( X(i,A) / X(A) ) / ( X(i) / X(world) )

An economy is roster-eligible at RCA ≥ τ (default 4) subject to two guards: it must hold at least $20M of the anchor basket, and its total merchandise exports must be at least $5B. Without the second guard the first run returned Gibraltar and Slovakia for EVs, Lebanon and Saint Kitts for transmission, and Tokelau and Wallis & Futuna, all re-export flags whose RCA is an artifact of a tiny denominator.

Large diversified exporters fall out at τ = 4 by construction, which is the point. Alternative 5 documents what happens when they do not.

5.3 Three measures

Each is normalised so that 1.0 means no association.

M1, roster capture. The share of a code's world exports held by roster economies, divided by the roster's share of world trade. Interpretable and directly readable off a table. Vulnerable to a roster economy being good at something unrelated.

M2, RCA-weighted lift. Continuous, with no roster cut. Every economy carries weight w(i) = min(RCA(i), 50) − 1 for RCA > 1 and zero otherwise. Numerator is the weighted sum of the code's export shares, denominator the weighted sum of baseline world shares. Robust to a single economy's unrelated specialisation, because a weight proportional to excess specialisation in the anchor basket does not transfer to whatever else that economy happens to export.

M3, proximity lift. P(RCA_code > 1 | RCA_anchor > 1) / P(RCA_code > 1), following Hausmann and Klinger's conditional-probability proximity, generalised from product pairs to product-versus-basket. Binary co-occurrence, discarding value.

5.4 The decision rule

Core status requires M2 ≥ 2.5. M1 ≥ 3.0 is a cross-check: where the two fall on opposite sides of their thresholds, the code goes to manual adjudication whichever way M2 fell. M3 is reported as a diagnostic and does not enter the rule.

M1 is deliberately denied veto power, and this is a correction to an earlier draft of this method. Section 8.2 shows that M1's ranking is largely a function of τ, its own roster cut: at τ = 2 wind retains 3 of its default top 10 and EVs retains 1 at τ = 6. A threshold on a scale that moves with a free parameter is a researcher degree of freedom wearing a measurement's clothes. M2 reads no roster cut and retains 10 of 10 across the entire sweep for every chain but biofuel.

M1 is nonetheless retained rather than dropped, for two reasons. Its margin on the benchmark true positives is 2.50x against M2's 1.22x, so it is the more legible measure to present. And a second view of a thin call is worth its cost: the 1.9% to 14.3% of codes on which the two disagree are precisely the cases where a screen should defer to a human.

M3 contributes no decision. Its benchmark margin is 1.04x, which is not a discriminating margin, and it is reported only because a measure that discards value and keeps only co-occurrence is a useful sanity check on the two that do not.

Three tiers result, and they are a label rather than a weight:

tiermeaningtreatment
coreheading is technology-specific and the screen agreesfull weight, sums into the headline
sharedgenuine chain trade in a heading that also carries large non-chain volumeretained, reported separately, never summed into the headline
excludedthe heading cannot contain the technology, or no associationdropped

A fractional split is deliberately not offered. A share of 0.5 on a code claimed by solar and wind is a number no one can source, and it renders both the chain figures and the total wrong at once. The present 84% inflation is at least a legible error.

5.5 Adjudication

Every code clearing the screen, plus every code above $5B of world trade regardless of score, is adjudicated against the HS heading's own legal text. The governing question is whether the heading can lawfully contain a good used in this technology in material volume. Chapter 87 headings cannot contain turbine parts. Ores and unwrought metals are inputs to everything and specific to nothing.

This step exists because of a documented failure. Unwrought zinc (790111) scores 5.84 on M1 for wind, above towers at 4.64, purely because Spain exports zinc. The screen produces a ranked shortlist of roughly 15 to 20 codes per chain, which replaces reading 89.

§6

Alternatives considered and rejected

All six regenerate via /usr/bin/python3 alternatives.py.

A1. Bill-of-materials value-addition share as the weight. A cost share is strictly positive for every component, so the rule cannot express "none of this line is this technology", which is the statement the dual-use codes require. It also undercuts the technology-specific lines by 4x to 20x: a tower is 26% of a turbine's cost and approximately 100% of world trade in HS 7308.20. BOM-weighting the seven-code wind list yields $13.55B against $35.49B on demand shares.

A2. The dictionary's own IO fields. direct_use_share and leontief_use_share down-weight raw materials correctly (copper ore 0.018 / 0.054). Rejected on the denominator: the fields are computed at the EXIOBASE-product level, so one value covers up to 21 distinct HS-6 memberships whose world trade differs by more than an order of magnitude, and summing across technologies for a single EXIOBASE product reaches 15.45 for "Chemicals nec". Whatever they are a share of, it is not a code's trade. They are also fully populated on Raw and Processed Material and nearly absent on Product Component (27%) and Process Equipment (2%), which is where a bill of materials lives.

Corrected 2026-08-19. This paragraph previously read that the fields "score motor-vehicle brakes identically to complete wind turbines, both 1.0 / 1.0", and that 58% of rows sit at 1.0. 1.0 is a missing marker, not a value. The rule "no EXIOBASE mapping implies never a fraction" holds with zero exceptions across all 715 rows, and 304 of the 350 rows at 1.0 carry a blank exiobase_product. Brakes and turbines agreed because the field had said nothing about either. The original verdict stands; the reason given for it did not, and on the computed subset the fields rank measured rates positively rather than failing to.

A3. Basket-specificity score. The share of a basket claimed by three or more chains, as a proxy for "this basket is generic". It ranks wind lowest at 24.8% when wind is the documented canonical offender, and batteries at 48.6% when batteries is documented clean.

A4. Naive exporter-profile similarity. Hellinger affinity between a candidate's export-share vector and the anchor's. Vehicle brakes score 0.735 against towers at 0.748 and miscellaneous plastics at 0.746. Every manufactured line is exported by the same large manufacturers, so raw profile similarity measures general manufacturing capacity.

A5. Roster capture including large diversified exporters. At RCA > 1 the wind roster is China, Germany, Denmark, Spain and India with a 26.55% baseline. Motor-generator parts score 1.73x against towers at 1.82x. At RCA ≥ 4 the roster is Denmark and Spain with a 2.24% baseline, and the same comparison is 1.86x against 4.64x.

A6. Within-basket RCA denominator. Rejected as endogenous: it measures specialisation within a universe the dictionary itself defines. Correcting it changes India's wind RCA from 4.2 to 3.1.

§7

Results

7.1 Anchor availability

Eight of eleven chains have a definitional anchor. Three do not, and that is a finding rather than a gap in execution.

chainanchor qualitywhy
wind, batteries, nuclear, heat pumps, magnets, biofuelcleanheading text names the technology
solarweakHS 8541.40 in HS12 merges photovoltaic cells with LEDs
transmissionweaktransformer codes are manufactured too widely to yield a discriminating roster
EVsnoneHS 8703.80 isolates battery-electric vehicles in HS17 and HS22. The station runs HS12, where they sit inside 8703.90 with hybrids and other cars
geothermalnoneno HS-6 line names geothermal; the basket is generic pumps, heat exchangers and drilling equipment
electrolyzersnoneHS 8543.30 covers electroplating as well as electrolysis

The EV result is a data problem for the station independent of this method. The dictionary carries HS22 descriptions against HS12 codes, so products.csv labels 870390 "vehicles with only electric motor for propulsion" while the $344B of trade recorded under it in HS12 is "other motor cars".

7.2 The screen, per chain

chainqualityanchorsanchor basketrosterbaselinecodes beforebeforecorecore $
batteriesclean1$118.5BHUN, POL2.17%70$838B7$135.2B
biofuelclean2$25.8BLVA, BEL, BGR, NLD4.66%74$314B11$29.4B
nuclearclean8$16.9BNAM, RUS, KAZ, GBR7.42%41$391B9$37.1B
windclean1$6.6BDNK, ESP2.24%89$1,036B5$25.5B
heat pumpsclean2$7.2BSWE, THA2.33%27$347B2$7.2B
magnetsclean2$6.7BPHL0.42%20$248B3$11.7B
solarweak1$76.0BLAO, KHM, VNM2.43%53$656B2$80.3B
transmissionweak2$4.8BMEX2.84%28$449B3$5.8B
electrolyzersnone062$467B
EVsnone099$1,613B
geothermalnone077$587B

Rosters are readable as industry facts, which is the first check on whether the construction is doing anything. Hungary and Poland are the EU battery manufacturing hubs. Namibia, Kazakhstan and Russia are uranium and enrichment. Denmark and Spain are wind. Mexico is transformers for the US market. Two are thin enough to state plainly: magnets rests on the Philippines alone at a 0.42% baseline, and transmission on Mexico alone.

Wind's screened basket is 5 codes and $25.5B against 89 codes and $1,036B, and its leaderboard becomes China, Germany, Denmark, India, Spain, the United States, Turkey, Italy, France and Portugal.

7.3 The cross-chain allocation problem largely dissolves

Section 3 separated membership from allocation. Resolving membership resolves most of allocation, because the codes that force a splitting rule are the codes that survive no chain's screen.

codes in 2+ chainsworld trade
before screening121$1,223B
after screening0$0B

Inflation falls from 84% to 0%. Copper ore's best score across its nine claimants is 1.33x and its worst is 0.00x, so it survives nowhere. The mechanism is structural: core status means concentration in a chain's specialists, and a line cannot be concentrated in Denmark and in Kazakhstan and in Hungary at once.

This result is computed over the eight screenable chains. The three without anchors are excluded from it, and they include the two largest baskets, so it should be read as a statement about the mechanism rather than a settled figure for the dictionary.

7.4 Adjudication

353 codes were adjudicated against HS heading text: 40 core, 56 shared, 257 excluded. Outcome by chain:

chainassertedcore codescore $share survivingverdict
nuclear$391.3B11$29.6B7.6%screen and heading text agree on the same 11 codes
wind$1,036.3B3$21.5B2.1%works; screen carries independent information
batteries$838.4B8$137.0B16.3%correct on legal scope; screen must not be used
biofuel$313.9B4$32.5B10.3%narrows; leaderboard is the industry
heat pumps$346.8B2$7.2B2.1%both subheadings name the technology
magnets$247.5B2$6.7B2.7%narrows, after the screen's top code is overruled
solar$656.3B1$76.0B11.6%narrows on legal scope, not on the screen
transmission$448.7B4$7.2B1.6%defect is membership, not breadth
EVs$1,612.8B0cannot be narrowed under HS12
geothermal$586.8B0no core tier; should carry no headline trade number
electrolyzers$467.1B0basket is a concordance output

Leaderboards after narrowing read as the industries. Nuclear becomes Kazakhstan, Russia, Canada, France, the UK and Namibia, which is the world fuel cycle. Magnets becomes China 61.4%, Japan 7.7%, Vietnam, the Philippines and Germany. Biofuel becomes US ethanol and pellets plus the Antwerp-Rotterdam-Germany biodiesel cluster. The asserted baskets had ranked Chile 1st and Peru 2nd in magnets, Australia 2nd in batteries, and Australia 1st in nuclear, all on ore.

Four findings that the screen could not have produced.

Transmission's problem is the opposite of breadth. Its transformer ladder carries 8504.31/32/33/34 and omits 8504.21/22/23, the liquid-dielectric transformers that are the actual traded grid product, $17.85B against $10.29B for the ladder it does carry. Those codes appear in no chain at all. Also absent: 8537.20 boards and panels above 1000 V, all of 8535 and 8536 (HV switchgear, breakers, isolators, surge arresters), and line insulators 8546.10 and 8546.20 while 8546.90 is present. Roughly $40B of the goods that define the chain are missing while it carries bottle-filling machinery.

Electrolyzers' basket is a concordance artifact. One HS92 code, 854380, walked forward into eight HS12 codes worth $160.7B, 34% of the basket, including $77.9B of ASML-class semiconductor manufacturing machinery at 8486.20. Chapter 84 Note 9 confines heading 8486 to semiconductor and display manufacture, so the membership is legally impossible rather than merely generous.

EVs is a vintage result with a measured magnitude. BACI HS17 splits what HS12 carries as 8703.90 into 870380 battery-electric ($139.75B) and 870340/50/60/70 hybrid and plug-in ($209.98B), summing to $352.53B against HS12's $344.36B, a 2.3% cross-vintage gap. The chain's single Final Product line is 39.6% battery-electric and 59.5% hybrid. Under HS12 the leaderboard is Germany, Japan, China; under the true BEV line Japan falls from 2nd to 5th and China rises from 12.6% to 23.4%.

A vintage collision double-counts by construction. The dictionary's HS22 rows 8501.61-64 (AC generators, explicitly other than photovoltaic) and 8501.80 (photovoltaic AC generators) are mutually exclusive in HS22, and the projector maps both onto the same HS12 lines. $7.32B is counted once into wind and once into solar on two readings of one code that cannot both be true. This belongs in the projector, not in the screen.

§8

Validation

8.1 Known-answer benchmark

Thirteen wind cases whose answer is settled by the HS nomenclature or by documented prior findings. Regenerates via /usr/bin/python3 benchmark.py, exit 1 on failure.

hs6truthM1M2M3
850231 complete turbinescore11.8411.1931.40
841290 rotor bladescore5.164.4212.08
850164 AC generators >750 kVAcore6.515.778.72
730820 towerscore4.642.938.72
850300 motor/generator partsnot core1.862.415.61
870830 vehicle brakesnot core1.661.508.37
732690 misc iron articlesnot core1.291.544.49
392690 misc plasticsnot core1.031.533.77
870840 vehicle gearboxesnot core1.011.274.19
281820 aluminanot core0.700.493.14
260300 copper orenot core0.520.090.00
854370 industrial elec. machinesnot core0.411.062.24
790111 unwrought zincnot core5.841.153.14

Separation of settled keeps from settled drops: M1 4.64 vs 1.86 (2.50x), M2 2.93 vs 2.41 (1.22x), M3 8.72 vs 8.37 (1.04x).

Under the rule in 5.4, 13 of 13 classify correctly: 12 decided automatically by M2, and one, zinc, routed to adjudication by an M1/M2 disagreement and excluded there on heading scope. M2 alone is sufficient on this benchmark.

The zinc row is the informative one. It is the case that separates the measures from one another, and it is the reason adjudication follows the screen rather than replacing it.

This benchmark caught two real errors during construction. An earlier draft asserted that zinc must score above 3.0 on M2 as a standing false positive; correcting the RCA denominator (A6) and adopting continuous weighting dropped it to 1.15, and the assertion failed. Separately, the first decision rule made core status conditional on M1, which Section 8.2 then showed to be unstable. Both findings were the fix.

8.2 Sensitivity

Each measure is swept against the parameters it actually reads. τ is the roster cut and enters M1 only. An earlier version of this test swept τ against an M2 ranking and returned perfect stability for all eleven chains, which was a check that could not fail.

M1, top-10 retained across τ ∈ {2, 3, 4, 6} × minimum economy size ∈ {$1B, $5B, $20B}. Ten of eleven chains fall below 8 of 10 somewhere on the grid. Wind retains 3 at τ = 2, EVs retains 1 at τ = 6, electrolyzers 0 at τ = 3, transmission 1 at τ = 6. Only nuclear is stable throughout.

M2, top-10 retained across minimum economy size ∈ {$0B, $1B, $5B, $20B, $50B}. Every chain retains 10 of 10 except biofuel, which falls to 7 at the $20B guard, and geothermal, which falls to 9 at $50B.

The asymmetry is the result. M1's threshold of 3.0 is meaningful only at τ = 4 and would have to be re-derived for any other cut. That is what disqualifies it as the deciding measure and what motivates the rule in 5.4.

8.3 Measure agreement

Pooled Spearman rank correlation across 640 (chain, code) observations:

pairρ
M1 vs M2+0.624
M1 vs M3+0.368
M2 vs M3+0.618

None approaches the ~0.95 at which the measures would be monotone transforms of one another and "require agreement" would be empty. They carry independent information.

Disagreement at the decision thresholds, which is the volume routed to adjudication:

chaincodes scoreddisagreerate
solar5311.9%
wind8944.5%
magnets2015.0%
heat pumps2727.4%
biofuel7479.5%
nuclear4149.8%
EVs991010.1%
transmission28310.7%
batteries70811.4%
geothermal771114.3%

Wind's four disagreements include 790111, the documented zinc case, which is the behaviour the routing rule is designed to produce.

8.4 Validity gates: when the screen is not measuring the chain

The adjudication pass established that the screen's validity is a property of the roster it derives rather than of the method. Two gates make that diagnosable before any adjudication, and both were derived from documented failures, so neither is a check that cannot fire. Run via /usr/bin/python3 validity.py.

Gate 1, the China-proxy test. Correlate M2 against China's share of each code across the chain's basket. Where the anchor's specialists are offshore assembly sites of a Chinese supply chain, the roster co-moves with China on every code and M2 stops encoding the chain.

chainrosterρ(M2, China's share)
batteriesHUN, POL+0.978china proxy
solarLAO, KHM, VNM+0.865china proxy
heat pumpsSWE, THA+0.573ok
magnetsPHL+0.479ok
transmissionMEX+0.429ok
windDNK, ESP+0.311ok
nuclearNAM, RUS, KAZ−0.286ok
biofuelLVA, BEL, BGR−0.328ok

Batteries fails in both directions, which is what a proxy does. Manganese metal, a steel and aluminium alloying input, scores 3.45. Lithium carbonate scores 0.10 because Chile ships 75% of it, and nickel sulphate 0.46 because Indonesia ships 40%. Keeping lithium hydroxide at 3.62 while dropping lithium carbonate at 0.10 would have been indefensible.

Gate 2, the sole-driver test, per code. Flag any code clearing the screen where one economy supplies 80% or more of its M2 lift.

An earlier per-chain version of this gate measured weight concentration and flagged wind at Denmark 76%, the one chain where the screen demonstrably works. Concentration is therefore not the discriminator. What separates Denmark in wind from the Philippines in magnets is whether a single economy carries one code's entire score. Denmark and Spain both contribute to towers and blades; only the Philippines carries nickel ore.

Retargeted per code, the gate returns zero flags on wind and catches the documented failures:

chainhs6M2share of liftdriver
magnets260400 nickel ore7.0199%PHL
batteries811100 manganese metal3.45100%CHN
batteries282520, 2805193.62, 2.50100%CHN
biofuel293292, 1510003.54, 2.5997%, 96%ESP
nuclear840110 reactors, 261210 uranium ore20.11, 16.36100%RUS, NAM

Nuclear's flags are the instructive case. Russia does export reactors and Namibia does export uranium ore, so those codes are correctly core. The gate routes to review rather than to rejection, and review confirmed them. A gate that only ever removed things would be a worse instrument.

Magnets' nickel ore is the strongest single argument in this exercise for keeping a human step. It scored M1 95.0 and M2 7.01, above both anchors, and would have been the chain's top-ranked code.

8.5 The null distribution, false-positive rate and discovery power

The first draft of this method reported no error rate of any kind. Every figure in Sections 7 and 8.1 to 8.4 is a statement about codes the dictionary already claimed, and none of them said how often the gate fires on a code with no relationship to the chain. An adversarial review identified this as fatal, and an independent reimplementation (null.py) reproduces its numbers.

Construction. BACI HS12 2024 holds 5,198 HS-6 lines. The eleven baskets claim 428 between them. The remaining 4,768 are, by the dictionary's own assertion, in no clean chain. Scoring those against each chain's roster gives M2 under "no relationship". Discovery power is then the number of non-anchor basket codes clearing the gate against the number expected if membership carried no information, matched on trade value within a factor of three, because M2 is not value-neutral.

chainnull median M2null p90FPR at 2.5non-anchor observedmatched expectedlift
nuclear0.601.855.3%51.63.22
biofuel0.902.298.2%105.71.76
transmission0.991.843.4%10.91.08
wind1.222.136.8%44.30.92
batteries1.052.6812.3%67.50.80
magnets1.012.8013.6%12.30.44
solar0.932.428.9%15.30.19
heat pumps1.191.984.2%00.90.00
total2828.50.98

Two conclusions follow, and neither is recoverable by moving the threshold.

A code clearing 2.5 is not thereby shown to belong anywhere. The gate fires on one in eight unrelated codes in batteries and one in seven in magnets.

The screen adds no detectable signal over chance. Wind, the chain where every other diagnostic said the method worked, returns 4 non-anchor survivors against 4.3 expected. The benchmark in 8.1 was passed on codes the anchors mechanically favour, which is why it read as a success.

The null also invalidates two earlier claims in this document. The stated null of 1.0 (Section 5.3) is a definition the data do not honour: it sits at 0.60 in nuclear and 1.22 in wind, so "no association" is half a unit apart depending on the chain. And the fixed 2.5 gate operates at the 86.4th percentile of the null in magnets and the 96.6th in transmission, an order-of-magnitude difference in stringency presented as one rule.

The 2.5 threshold itself was set by a single observation, lowered from 3.0 so that wind towers at 2.93 would survive. The 15 codes that relaxation admits arrive against 13.2 expected by chance, a ratio of 1.14.

8.6 Other findings from adversarial review

Three lenses ran against the method: statistical, trade-classification and reproducibility. Findings not already folded into the sections above, all with reproduction evidence in out/verification.json.

The benchmark was decoupled from the artifact it validated. benchmark.py hardcoded rca_vector(["850231"]) and never read anchors.json. Seeding the shipped wind anchor to copper ore inverted the entire wind result (850231 falls 11.19 to 0.008, copper ore rises 0.087 to 41.3, and the wind core becomes eight ore and metal codes) while the benchmark stayed green. Fixed, and the defect was re-seeded to confirm the fix reports six named failures rather than passing or crashing.

Effective sample size. M2 is frequently a one-country statistic presented as a weighted average. 43.5% of all 402 scored codes have an effective number of contributing economies below 2, including all seven battery survivors and all three magnet survivors. Wind's own anchor draws 77.8% of its numerator from Denmark.

Denominator choice moves outcomes. Recomputing RCA on manufactures only (HS 28-97) rather than total merchandise changes the pass set in five of eight chains, including nuclear (9 to 12 codes) and batteries (7 to 3).

No materiality floor on the scored code. The $20M floor applies to roster eligibility only, so a $1M line can top a chain: EVs' 846241 scores M2 4.29 on $0.001B. Three separately reported false positives trace to this one mechanism.

A legal argument in the benchmark may be inverted. The benchmark treats 841290 as a wind true positive and 850300 as a true negative. Section XVI Note 2(b) classifies parts suitable for use solely or principally with machines of heading 85.01/85.02 in 8503.00; a grid turbine is 8502.31, so its blades and nacelle parts arguably fall to 8503.00 rather than 8412.90. Denmark's exports of the two codes are equal to two decimal places ($0.58B each). This is unresolved and is flagged rather than fixed, because resolving it requires a WCO classification ruling rather than more data.

Three membership holes found while adjudicating. Solar contains no inverter code: HS 8504.40, $94.53B, the largest non-module component of a PV system, is claimed by EVs alone. Heat pumps carries the undifferentiated parts subheading 8415.90 ($30.51B) while omitting four of the five machine subheadings of heading 8415 ($34.58B). Transmission omits heading 8535 entirely ($12.18B of HV switchgear).

Namibia's nuclear roster membership rests on a misdeclaration. Its entire anchor exposure is $97.9M of HS 2844.20, enriched uranium, which Namibia has no capacity to produce. Its actual exports are U3O8 (2844.10) and uranium ore (2612.10), neither of which is an anchor code.

§9

Limitations

The method requires a definitional anchor, and three chains do not have one. EVs, geothermal and electrolyzers carry no HS-6 line that isolates the technology. The EV case is a vintage artifact worth reporting on its own: HS 8703.80 (vehicles with only electric motor for propulsion) exists in HS17 and HS22, and the station runs HS12, where those vehicles fall inside 8703.90 alongside hybrids and other cars. The dictionary carries HS22 descriptions against HS12 codes, so products.csv labels 870390 as the electric-vehicle line while the trade under it is not. Geothermal has no HS line that names it, and 854330 covers electroplating as well as electrolysis. For these chains the screen has no power and every code requires direct adjudication.

Anchor quality bounds everything downstream. A polluted anchor produces a polluted roster and every score inherits it. Solar and transmission are marked weak for this reason: 854140 in HS12 merges photovoltaic cells with LEDs, and the transformer codes are made in too many places to yield a discriminating roster.

The method requires geographic concentration, which is a property of the industry. Wind works because Denmark exports 17.4% of world turbines on 0.4% of world trade. Transmission has no equivalent, and its screened basket comes out at 3 codes and $5.8B, which reflects an absent roster rather than a clean basket. The screen is a triage device that reduces 89 codes to roughly 15 for reading. Where triage fails, the fallback is to adjudicate the whole basket, which for transmission is 28 codes and entirely tractable.

The anchor set is drawn from the dictionary under test. If the dictionary mis-assigns a Final Product row, the chain's screen is wrong in a way the screen cannot detect. Manual anchor validation mitigates this and does not eliminate it. An independent anchor list, authored from HS nomenclature without reference to the dictionary, would close it.

Association is not content. The screen measures whether a line's exporters look like a technology's specialists. Unwrought zinc scores 5.84 on M1 for wind because Spain exports zinc. Adjudication against heading scope is the only step that examines what the good actually is, and it is a human step.

One year, one side. Everything is 2024 exports. Whether a code's association is stable across years is untested, and import-side specialisation is unused. Entrepôt and re-export economies distort Balassa RCA; the $5B size guard removes Gibraltar, Tokelau and Saint Kitts and does not address Hong Kong, Singapore or the Netherlands.

Membership is solved; attribution is bounded, not solved. A core code still enters at weight 1.0. Some share of HS 7308.20 is telecommunications and transmission masts rather than wind towers, and this method does not say how much. Section 3's question (2) remains open, and the shared tier exists to keep the honest cases out of the headline rather than to price them.

The thresholds are conventions. M2 ≥ 2.5 and M1 ≥ 3.0 are calibrated on a thirteen-case wind benchmark. Section 8.2 shows M2's ranking is stable across the economy-size guard, which is a different claim from the threshold being right. A larger benchmark spanning more chains would test the threshold rather than the ranking, and it does not yet exist.

§10

Reproduction

``bash cd story-stations/clean-trade-pulse/methodology ./reproduce.sh ``

anaconda's numpy is broken machine-wide on the authoring host; every script targets /usr/bin/python3. Step 1 reads 11.1M BACI records and takes roughly 35 seconds. All other steps are seconds.

scriptproduces
scan_totals.pyout/baci_2024_totals.csv, the Balassa denominator
engine.pyout/screen.json, M1/M2/M3 per (chain, code)
alternatives.pySection 6, all six rejected rules
benchmark.pySection 8.1, exit 1 on failure
tables.pyevery numbered table
validate.pySections 8.2 and 8.3, sensitivity and measure agreement
validity.pySection 8.4, the two validity gates
null.pySection 8.5, null distribution and discovery power
appendix.pySection 11, full per-chain listings
§11

Appendix

Full per-chain code listings, every score and every decision, regenerate via /usr/bin/python3 appendix.py > out/appendix.md. Adjudications with rationale are in out/adjudications.json (353 rows), verification findings in out/verification.json.

§12

What to do

Adjudicate. Do not screen. The narrowing that worked was 353 readings of HS heading text against the question "can this heading lawfully contain this technology, in material volume". That procedure is reproducible, auditable one row at a time, and it produced every result in Section 7. It cost roughly 20 to 30 codes of close reading per chain, which is a day's work across all eleven, and there is no evidence the statistical screen made it cheaper.

Keep the tier model. core / shared / excluded as a label rather than a weight. Headline figures sum core only; shared is retained and reported separately. The concrete case for never summing them: biofuel's core leaderboard is US ethanol and the Antwerp-Rotterdam biodiesel cluster, and adding the shared tier pulls Trinidad to 3.0% and Saudi Arabia to 2.6% on merchant methanol.

Do not adopt a fractional split. A share of 0.5 on a code claimed by solar and wind is a number nobody can source, and it makes the chain figures and the total wrong at once. Section 7.3's result stands on the adjudicated data: after adjudication the pairwise overlap between wind, solar and batteries is exactly zero codes, because the shared lines were residual baskets and ores that fail in every chain that claimed them.

Keep the two validity gates (8.4) as diagnostics. They correctly identify batteries and solar as chains where any roster-based measure degenerates into a China proxy, and they flag single-economy codes for review. They are cheap and they fire on real cases.

Fix these in the projector, not the dictionary, and not by weighting.

  1. The HS22 8501.61-64 / 8501.80 collision, which double-counts $7.32B into wind and solar on two mutually exclusive readings of one HS12 line.
  2. The HS92 854380 walk-forward that puts $160.7B of semiconductor and display manufacturing equipment into electrolyzers, contrary to Chapter 84 Note 9.
  3. HS12 descriptions carrying HS22 semantics, which is how products.csv labels 870390 as the battery-electric vehicle line when the HS12 trade under it is 39.6% BEV and 59.5% hybrid.

Fix these in the dictionary, upstream at ecoclassical/NZIPL-CVCE.

  1. Transmission's membership hole: add 8504.21/22/23 (liquid-dielectric transformers, $17.85B), 8537.20, heading 8535 ($12.18B), and line insulators 8546.10 / 8546.20.
  2. Solar has no inverter code; 8504.40 ($94.53B) is claimed by EVs alone.
  3. Heat pumps carries 8415.90 parts while omitting four of five 8415 machine subheadings.

Three chains should not carry a headline trade number as things stand. Geothermal has no HS line that names it. Electrolyzers' only candidate heading also covers electroplating. EVs cannot be separated from hybrids under HS12, and that one is fixable by moving the station to HS17 rather than by any weighting scheme.

If the screen is revisited, the minimum bar is: an anchor list authored from the nomenclature without reference to the dictionary, a per-chain null with a stated false-positive rate, a threshold calibrated as a null percentile rather than a fixed number, and a held-out chain the method has never seen. Absent those, it should be described as a presentation ordering and not as evidence.