Skip to content
Heterodata An Arcanum Research project Shaikh Data
Shaikh Data

Shaikh Data

Open, reproducible data behind Anwar Shaikh's economics — built with the Anu framework

A hub of fully-documented, reproducible economic datasets reconstructing the empirical material in Anwar Shaikh's work. Every series is traceable to a named source, a construction method, and a provenance record, and is explorable as an interactive chart. Built with the Anu data-construction pipeline; published openly.

Programmatic access

Every series sits at a stable download URL (CSV, XLSX, and Parquet — no account, no API key), and /llms.txt gives your agent a machine-readable map of the data, the code, and the outputs. A Robin R / Python client package for pulling series by ID is planned; the downloads above are the supported interface today.

Datasets

Measuring the Wealth of Nations

Replication and 1990–2025 extension of Shaikh & Tonak (1994)

An open replication and modern extension of Shaikh & Tonak's Measuring the Wealth of Nations: The Political Economy of National Accounts (Cambridge, 1994). 59 published Marxian national-accounting series — total product, constant and variable capital, surplus value, the rate of exploitation, the profit rate and the net social wage — for the book period (1948–1989) plus an extension built from BEA and BLS data, alongside replications of the follow-up studies the project actually carries: Tonak (1984), Shaikh & Tonak (1987), Shaikh & Tonak (2002), Moos (2017), Mohun (2005), Mohun (2014) and Cronin (2001). Every series and subseries resolves to a full bibliographic record.

59 series 1948–2025 33 primary · 26 extra

Shaikh Capitalism Data

Replication and 1860–2025 extension of Shaikh (2016) Capitalism

An open replication and historical extension of the empirical material in Anwar Shaikh's Capitalism: Competition, Conflict, Crises (Oxford, 2016). 110 published series across the book's chapters, the book appendix, and the Sraffa-price follow-up studies — each chart traceable to a named source, a construction method, a provenance record and a full bibliographic citation.

110 series 1860–2025 95 primary · 15 extra

The comparative layer

The competing measurements of the same categories, side by side

Where scholars have measured the same Marxian categories from the same US accounts and reached different numbers, this layer puts those measurements beside one another instead of letting one stand in for the rest. It carries the net social wage as four separate lines (Shaikh & Tonak's book line, the 2000 revision, Tonak's 2025 update and Moos's 2019 replication), Mohun's own published recomputation of the classification, and the source-vintage dossiers that explain why the underlying BEA data moved. Every value is a digitization of what its author actually printed or stored, attributed to that author and resolved to a full bibliographic record.

136 series 136 primary · 0 extra

Built with the Anu framework

Every dataset here is constructed by the same staged, documented pipeline. The framework uses two unrelated letter-prefix vocabularies — keeping them distinct is what makes the build honest.

Pipeline phases — what a script does

The build is a sequence of phases, each with its own script prefix: L Loading (fetch + unit-check raw source data), P Processing (construction and transformation, and processing only), V Validation (check outputs against the published benchmarks), M Manual adjustment (documented hand corrections), A Analysis, and O Output (the tidy CSVs, Parquet, and provenance records you download here). A script's prefix tells you what it does — nothing about how a series is classified.

Series IDs — what a series is

Separately, every published series has an ID. S### is a primary series — a figure or table from the book or study being replicated (the digits encode chapter and sequence). XS is an extra series — a book-appendix series or a series drawn from another study, always listed after the primary series. A series ID's first letter says nothing about pipeline stages.