Week 10 · International Institutions

International institutions: instrument or authority?

An applied debate. One camp, rational-design institutionalism, reads institutions as the self-conscious instruments of states, designed dimension by dimension to solve cooperation problems. The other, Barnett & Finnemore, reads International Organizations as autonomous, constructed authorities: bureaucracies that make rules, define reality, and drift beyond what any state asked for. The week tests both lenses against the UN Security Council, R2P, and the ICC.

Level III: International Institutions

Paradigm: Applied Debate · Rational Design vs. IO Authority III

The clash here is not realism versus liberalism but two ways of seeing the same building. For Koremenos, Lipson & Snidal, an institution’s shape, including who is in, what it covers, how centralized it is, who controls it, and how flexible it is, is the rational output of states solving distribution, enforcement, and uncertainty problems. For Barnett & Finnemore, the same institution is a Weberian bureaucracy whose rule-making authority lets it constitute the world, exercise power states never delegated, and creep beyond its mandate. Beardsley & Schmidt, Hehir, and Jo & Simmons then become referee cases: does the UN follow the flag or the Charter, is R2P a real norm, can the ICC deter?

Reading list

  • Koremenos et al. 2001. “The Rational Design of International Institutions.” International Organization.
  • Barnett & Finnemore. 2004. Rules of the World, Ch. 1 & 4.
  • Beardsley & Schmidt. 2012. “Following the Flag or Following the Charter?” International Studies Quarterly.
  • Hehir. 2013. “The Permanence of Inconsistency: Libya, the Security Council, and the Responsibility to Protect.” International Security.
  • Jo & Simmons. 2016. “Can the International Criminal Court Deter Atrocity?” International Organization.
  • Application: Ovadek. 2025. “Hungary’s Exit from the International Criminal Court is a Sign of the Times.” The Conversation.

Koremenos et al. (2001): The Rational Design of International Institutions

International Organization · institutions as states’ purposeful, problem-solving designs

They define international institutions as explicit arrangements, negotiated among international actors, that prescribe, proscribe, and/or authorize behavior. The arrangements are explicit (public, at least among the parties) and are the fruits of agreement. This covers formal organizations with a bureaucracy and budget, the WHO, the ILO, but also well-defined informal arrangements like “diplomatic immunity” that have no formal bureaucracy or enforcement mechanism yet are fundamental to the conduct of international affairs. So “institution” is broader than “organization”: the common thread is an explicit, negotiated rule, not a building full of staff. The core claim is that design differences are not random: they result from rational, purposive interaction among states and other actors solving specific problems. States, acting for self-interested reasons, design institutions deliberately to advance their joint interests. Against constructivism: they don’t deny that institutions spread norms, “normative discourse is an important aspect of institutional life”, but they downgrade norm-diffusion as not the main engine of design, and reject treating IOs as purely exogenous outside forces, since institutions are the self-conscious creation of states (and, secondarily, of interest groups and corporations). Against realism: their move is the exact opposite. Realists treat institutions as little more than ciphers for state power; Koremenos et al. say that exaggerates a real point. States rarely let institutions become major autonomous actors, but institutions are far more than empty vessels; states spend serious time and effort building them precisely because they can advance or impede state goals in economics, the environment, and security. The five dimensions are Membership (who is in), Scope (what issues are covered), Centralization (whether tasks such as information, bargaining, dispute settlement, and sometimes enforcement run through a focal body), Control (decision rules: voting, vetoes, weighted votes), and Flexibility (escape clauses, renegotiation, sunset provisions, and other ways to adapt). The conjectures then rest on four assumptions: (1) rational design, meaning states design purposefully to advance joint interests; (2) the shadow of the future, meaning future gains are valued enough to sustain cooperation; (3) transaction costs, meaning building and running institutions is costly; and (4) risk aversion, meaning states worry about adverse effects when creating or modifying institutions. I find the dimensions clearly stated and the assumptions broadly sensible, though they presuppose a deliberate, well-informed designer more than messy real bargaining often is.

The shadow of the future (Axelrod & Keohane 1985) is the idea that repeated interaction makes states value future consequences, so they defect less today for fear of retaliation tomorrow; institutions can lengthen that shadow, alter payoff structures, and manage multi-actor games, enabling cooperation even under anarchy. Two cooperation problems then need to be kept apart. A distribution problem exists when everyone agrees cooperation is good but disagrees on which cooperative outcome to pick, bargaining over the surplus, who gets what share; nobody wants to defect from cooperation itself, they want a better split (e.g., tariff-cut formulas in GATT/WTO, or emission-quota allocation in climate, where the US, EU, China, and the G-77 all want a deal but different cuts). An enforcement problem exists when, even after we pick the same outcome, each side is tempted to cheat ex post for an immediate gain, measured by the minimum discount factor needed to sustain cooperation (e.g., hidden export subsidies or import barriers after signing WTO commitments; OPEC members signing quotas, then overproducing). In short: distribution asks which cooperative point we pick; enforcement asks whether we’ll stick to it. My main revision concerns path dependence and stickiness. They mention it (the UN Security Council example) but don’t model how initial control rules, once locked in, can later be exploited to paralyze the institution, exactly what the US does by blocking Appellate Body appointments at the WTO. They note (p. 773) that control and centralization can be varied independently, which is a real insight; but the corollary they miss is that one actor can then freeze the centralized part without formally defecting. So I’d push the framework to treat control rules as endogenous weapons over time, not just as design choices at founding.

Barnett & Finnemore (2004): Rules of the World

Ch. 1 & 4 · IOs as bureaucracies, authority, autonomy, and pathology

The functionalist account says IOs are created and survive because of the (usually desirable) functions they perform: states build them to solve problems of incomplete information, transaction costs, and other barriers to welfare improvement. Their alternative is constructivist: it treats IOs as autonomous actors and explains the power they use, their tendency toward dysfunctional and even pathological behavior, and how they change over time. The foundation is that IOs are bureaucracies, a distinctive social form of authority with its own internal logic and behavioral tendencies. Bureaucracies exercise power through the ability to make impersonal rules, and they then use those rules not just to regulate but to constitute and construct the social world: creating new categories of actors, forming new interests, defining shared international tasks, and diffusing models of social organization globally. But the very impersonal rules that make bureaucracies effective can also cause problems. So: functionalism sees IOs as tools that fix cooperation problems; their alternative sees IOs as rule-making, authority-claiming bureaucracies whose internal culture and authority can make them do things states never asked for. Principal-agent analysis treats IOs as agents of states (principals). It recognizes that states purposely grant some autonomy, because otherwise the IO couldn’t perform its tasks; within a “zone of discretion” the IO can advance state interests, make policy where state preferences are weak or unclear, and sometimes even act against some states’ interests, driven by the gap between what agents and principals want. Its strengths are real: it sees that delegation requires discretion. But it still treats autonomy as residual or accidental, something to be reined in by monitoring and sanctions, and assumes the agent’s preferences are basically derived from the principal. IOs diverge when (1) their authority source pushes them: the IMF’s expert authority keeps it pushing macro-stabilization even where its Articles didn’t clearly allow it; UNHCR’s moral-plus-expert authority kept expanding who counts as a refugee even when states weren’t demanding it. (2) Path-dependent rules tell staff “this is how we solve the problem” even when states are indifferent, IOs can operate where states are indifferent and even oppose or change state interests. (3) Field offices diverge from capitals: UNHCR field staff sometimes wanted to protect refugees more than states did and resisted forced return, while later the internalized “repatriation” script let UNHCR push refugees back even when host states would have left them in camps, bureaucratic, not statist. More broadly, IOs clash with states when their universal solutions (climate mitigation, trade fairness) cut against a state’s protection of domestic industry; when their attention to human rights or governance is read as interference in internal affairs; when sovereignty and the claim that domestic law trumps international law collide; when reform threatens the existing power balance; and when universalist values (democracy, liberty) clash with a state’s political system, culture, or beliefs. Both schools focus on only two tools of power, material coercion/inducement and information, and only on the IO’s ability to shape state behavior, which barely captures how IOs shape the world around them. For Barnett & Finnemore, IOs are powerful not mainly because they hold material and informational resources but because they use authority to orient action and create social reality: they don’t just move information, they interpret it, investing it with meaning and turning information into knowledge. The World Bank doesn’t only gather data, it defines what development is, what measures it, what counts as poverty, and what to do about it. As authorities, IOs exercise power two ways: regulatory power, altering incentives so actors conform to existing rules (the UN Human Rights Commission publishing torture data creates compliance incentives); and constitutive / social-construction power, deciding whether there is a problem at all, what kind it is, and whose responsibility it is, UNHCR helps define what a refugee is, the IMF and Bank define “best practices” and “good governance,” and human-rights bodies define what rights are. A third channel is power through rules that stick even when wrong, since bureaucracies become obsessed with their own procedures (their “pathologies”). What the neo-neo approaches miss is therefore constitutive power, the possibility that IOs fail in ways states can’t be blamed for, and IO change that states neither ask for nor anticipate. I largely agree: the constitutive, agenda-setting power of bodies like the Bank or UNHCR is real and invisible to a purely material or informational account.

Treating IOs as bureaucracies starts from Weber: this means rational-legal authority with specialization, rules, and expert knowledge. The leverage is threefold. First, autonomy is explained as flowing from authority, not simply from lax state control, “it is because of their authority that bureaucracies have autonomy and the ability to change the world.” Second, pathologies are seen as organically produced by the same features that make IOs useful: division of labor breeds tunnel vision; standardized rules breed inflexibility. Third, change is traceable as the path-dependent encoding of experience into governing rules. On failure: sometimes it is clearly the states’ fault, staff are handed mandates without funding and tasks others won’t do. But not all undesirable behavior is states’ fault; IOs generate their own mistakes, perversities, even disasters, and it is often the very features that make them authoritative that breed dysfunction. “Mission creep” captures this: missions may expand because states add tasks, but there is an unintended internal logic too, as IOs define problems and implement mandates, they do so in ways that permit or require more intervention by more IOs. Examples from the book: the IMF moving from balance-of-payments financing to full-scale domestic economic reform; UNHCR from 1951/Europe/persecution to global, mixed-cause, material assistance, and finally to “solving” displacement via repatriation; the UN Secretariat from interstate peacekeeping to peacebuilding, state-making, and civilian policing. The cause is Weberian-constructivist: rational-legal authority likes to universalize its categories, so once you build a rule (“refugee problems should be solved; the best solution is going home”) you keep meeting cases that “fit” it, and you creep. Which features matter? Strong, widely accepted expert or moral authority plus looser day-to-day oversight make an IO more likely to act autonomously; the ability to define categories and to convert information into knowledge makes it more powerful vis-à-vis states; and heavy rule-boundedness with high specialization, or dependence on “passing the hat” for funds, makes it more likely to fail or turn pathological. UNHCR is the UN High Commissioner for Refugees, mandated to protect refugees, the displaced, and the stateless. The Ch. 4 takeaways: the early-Cold-War, decolonization, and Asian crises exposed the 1951 Convention mandate as too European, too narrow, too political; UNHCR found a procedural, low-visibility, morally framed workaround, “good offices”, that redefined who could be helped (no need to prove persecution, no need to be in Europe). Over time this hardened into a “repatriation culture” in which the best solution to displacement was going home, which in practice produced involuntary or semi-voluntary returns (e.g., Burma/Rohingya), violating non-refoulement, the very rights classical refugee law protects. It is a proof by hard case: even an IO that is financially dependent, operating in a politicized area, and facing clear state preferences still shows organizational effects. It fits the theory well, the invention of “good offices” is exactly change “states do not ask for or anticipate”; the moral-plus-expert authority mix is why they chose it; the repatriation pathology shows impersonal rules causing harm. But there are robust alternatives the authors concede make these hard cases: a pure state-pressure story (Britain wanting Hong Kong stabilized, Asian governments wanting camps emptied, Western donors opposing permanent settlement); a resource-dependence story (because UNHCR passes the hat, it tailors policy to donors); and a norm-cascade story (post-Cold War return operations in Cambodia, Mozambique, and Central America made “return” the norm, and UNHCR merely carried it). B&F reply that these pressures got encoded into UNHCR’s own rules, which is what makes the effect bureaucratic rather than merely responsive. On quality: strengths are the hard-case design, the process-tracing logic, and a clean match between the key variable (type of authority) and case choice (IMF = expert, Secretariat = moral, UNHCR = mix). Limits: they infer a “culture” from a small number of policy episodes (mainly 1980s–90s shifts and the Rohingya case), which risks over-reading and cherry-picking; the heavy financial dependence means a good rationalist could explain much with donor and host-state pressure; and there is no counterfactual, what if leadership had been more legalist/protection-oriented? The text itself shows leadership mattered enormously (pragmatist High Commissioners like Hocke, Stoltenberg, and Ogata pushing repatriation, while fundamentalists clustered in the legal and protection divisions), which raises the deep question of whether “culture” is cause or effect. The case shows beautifully how culture works, but doesn’t fully prove that culture is the main driver, rationalist and constructivist readings remain compatible with the evidence, so it needs more cases and tighter controls. It depends on the research question, and the two are best read as a division of labor across the institution’s life. Koremenos, Lipson & Snidal start from states’ problems (distribution, enforcement, uncertainty, number, asymmetry) and predict variation in features (membership, scope, centralization, control, flexibility): IO form is a function of state preferences and problems. Barnett & Finnemore start from the IO as a bureaucracy and explain behavior after creation: autonomy, mission creep, pathology, authority-based expansion. If the question is “why was this IO designed this way in 1951?”, Koremenos et al. gives the sharper, more parsimonious answer. If the question is “why, by the 1990s, is this IO doing things no state ever wrote into its charter?”, Barnett & Finnemore is the only one that actually looks inside the organization. For the UNHCR chapter specifically, the Hong Kong “good offices” episode and the later repatriation culture, the puzzle is precisely expansion beyond mandate, which a rational-design story can only label “drift” or “noncompliance.” So on this evidence I lean B&F, while granting that design rationalism remains the better account of founding moments.

Beardsley & Schmidt (2012): Following the Flag or Following the Charter?

International Studies Quarterly · does the UN intervene for its mission or for P-5 interests?

Why does the UN intervene more strongly in some conflicts than others? Two answers compete: (a) how much a crisis threatens international peace and stability, the UN’s “organizational mission”; or (b) how much it touches the private, parochial interests of the five veto-using Security Council members (the P-5). The question speaks straight to the IO-autonomy-vs-great-power-control debate. If UN action tracks human suffering and escalation, that supports the Barnett & Finnemore / Finnemore (2009) view that IOs can act on their mandate; if it tracks big-power strategic interests, it supports the realist complaint that IOs are mere vehicles for the parochial interests of their most powerful members. The parochial-interest model says UN involvement is simply a function of the P-5’s private interests; the organizational-mission model says it primarily reflects how far a conflict challenges the UN’s Charter mandate of promoting international peace and stability. The difference is whose problem the UN is solving: the P-5’s private problem or the Charter’s collective one. Structurally, the UN (founded 1945, 193 members, six main organs) decides mainly through a universally representative General Assembly whose resolutions are only advisory, and a 15-member Security Council whose Chapter VII resolutions are legally binding: non-procedural votes need nine in favor including all five permanent members (China, France, Russia, UK, US), so any P-5 veto blocks action. Only the Council can determine a threat to peace and authorize force (Article 42), ranging from consent-based traditional peacekeeping to “robust” peacekeeping (“all necessary means”) to enforcement without consent. Through a Koremenos lens the UN is almost a “trapped” institution: near-universal membership aids distribution and legitimacy but creates huge enforcement problems (M1 vs M3 tension); P-5 veto concentrates control but produces paralysis (it inverts V1, where individual control should fall as number rises); centralization is high for information but weak for enforcement because sovereignty blocks C4; scope is vast (security, development, rights, environment, health), aiding cross-issue linkage but diluting focus; the Charter is nearly unamendable (2/3 of the Assembly plus 2/3 ratification including all P-5), so flexibility comes only through creative interpretation, contradicting F1’s call for flexibility under uncertainty; and Council composition no longer tracks contribution (Germany, Japan, Africa, Latin America under-represented), violating V2. The synthesis: these aren’t failures of rationality but rational design under extreme constraints: sovereignty, great-power politics, and distributional conflicts that make reform impossible because members benefit differently from the status quo. The UN is a second-best solution, not an optimal one.

The sample is the International Crisis Behavior dataset (v7, 1945–2002); the unit is the crisis (N = 272). The dependent variable is a 7-point level of UN involvement; because it is ordinal, they use ordered probit. In the parochial model (Model 3, full sample), a defense pact with a P-5 raises involvement (0.525, the P-5 wants cover/burden-sharing); “P-5 against non-P-5” lowers it (−0.439, when a great power fights a small one, the UN tends to stay out); “post-Cold War” raises overall activity (0.941); and the key interaction, “P-5 against non-P-5 × post-Cold War” (1.378), flips the earlier negative effect positive: once the P-5 cooperate more, they will use the UN even in their own crises. This is their sharpest test: salience matters only when P-5 overlap allows it. The defense-pact effect is itself eaten away after the Cold War (−0.873), suggesting allies’ crises no longer automatically pull in the UN once the P-5 prefer to act unilaterally or through other channels. In the organizational-mission model (Model 4), all four mission variables are clean and significant: level of violence (0.281), number of crisis actors (0.233), and logged duration (0.128) all raise involvement, while contiguity (−0.417) lowers it (border skirmishes look less likely to spill over). Model 4 beats Model 3 on AIC/BIC and proportional reduction in error (2.76 vs 7.87), and wins a Clarke non-nested test at .001; the combined Model 5 keeps all four mission variables stable while the P-5 variables thin out (“P-5 against P-5” becomes −0.774, confirming that two great powers on opposite sides paralyze the UN), implying that some P-5 variables were really proxying crisis complexity, not parochial interest. Figure 2 visualizes the substantive effects: mission variables on the right are large (number of crisis actors ≈ +4.67 for operational deployments; post-Cold War ≈ +1.78; violence ≈ +1.23), most P-5 variables on the left are small or negative, and the sign-flip is clearly displayed. It makes the point that mission beats P-5, far more clearly than raw probit coefficients. I’d change one thing: the long, crowded labels should be grouped into three blocks (Cold-War P-5, post-Cold-War P-5, mission) so the reader instantly sees P-5-alone = weak, P-5 × cooperation = stronger, mission = strongest. Broadly yes, the theory is clear, matches the data, and the mission model genuinely outperforms. But I’d note the hypotheses don’t break out issue-area crises (resource/maritime crises may draw UN attention for reasons other than violence), and a clean observable implication is missing: if the UN really follows the Charter, regional spillover crises (e.g., Arab–Israeli 1967 → 1973) should keep drawing attention even when no P-5 has a fresh stake, a nice case-level test. Two things I’d still want, one of which they concede: better timing data, and a disaggregation of Security-Council-authorized versus Secretariat-driven action, since some of the coded activity “does not represent any UN Security Council input.” That is the one spot where the P-5 model could look artificially weak.

Hehir (2013): The Permanence of Inconsistency

International Security · Libya, the Security Council, and the Responsibility to Protect

The Responsibility to Protect was born after Rwanda and Bosnia exposed the moral failure of sovereignty-as-non-interference. The 2001 ICISS report redefined sovereignty as responsibility, states must protect their populations from mass atrocities, and if they fail the international community must act (the emergence stage, led by norm entrepreneurs: Canada, scholars, NGOs). The 2005 World Summit adopted R2P; 2009 codified three pillars (the state’s duty, the international community’s duty to assist, and the Security Council’s authority to act collectively when a state fails), a tipping point. In 2011 the Council explicitly invoked R2P over Libya, authorizing “all necessary measures”, R2P’s high-water mark and apparent cascade. But NATO’s intervention slid into regime change, generating backlash; Syria’s atrocities ran into Russian and Chinese vetoes; the Global South grew suspicious that R2P masks Western-led regime change (Brazil’s “Responsibility while Protecting” stressed proportionality and accountability). The life cycle stalled before internalization; the cascade was interrupted. Hehir’s verdict is skeptical but not hostile: protecting civilians is a good moral principle, but R2P is not yet a norm in the strict sense. It has not been internalized or applied consistently. It is a political idea, not a legal rule; states invoke or ignore it by interest, so intervention decisions “will continue to be made in an ad hoc fashion.” R2P is rhetoric and a moral vocabulary, a “loud voice in a large, disparate, chanting crowd”, not a stable guide to behavior. I find this persuasive: invocation without consistent application is exactly what a slogan, not a norm, looks like. He gives a multi-cause explanation and calls the intervention an exception produced by coincidence, not the rise of a new global norm. The overlapping factors: Qaddafi’s shocking threats to massacre Benghazi created media and moral pressure; regional support, especially from the Arab League, made action politically easier for Western and P-5 states; the abstentions of China and Russia, swayed by regional opinion, removed the veto obstacle; domestic politics and personalities mattered (Sarkozy’s ambition, Obama’s cautious pragmatism, Cameron’s moral framing); and Libya’s strategic location plus Qaddafi’s pariah status helped. He sums it up as “a unique constellation of necessarily temporal factors”, a one-time mix that won’t easily repeat. So Libya is an aberration, not a precedent: not proof that R2P became law or norm, not a Western oil conspiracy, but a temporary alignment of humanitarian concern and national interest that enabled quick action. Once those conditions vanished, as in Syria soon after, the inconsistency returned. Hence the title: “The Permanence of Inconsistency.”

After Resolution 1973, an “R2P triumph” chorus emerged. Their evidence had four strands: the action matched R2P’s text (authorizing “all necessary measures to protect civilians”; Gareth Evans called it “a textbook case of the R2P norm working,” Ban Ki-moon said it affirmed the international community’s determination to protect); multilateral and regional backing (Arab League and African Union support, read as norm internalization beyond the West); humanitarian urgency and moral legitimacy (Qaddafi’s massacre threats as a classic R2P trigger); and elite optimism (Axworthy’s “dawn of a more humane world,” Bellamy & Williams’s “new politics of protection”). Hehir doesn’t deny the intervention happened; he attacks the inference. First, they mistake outcome-similarity for causal proof: the two resolutions (1970, 1973) mention only the “Libyan authorities” responsibility to protect their population’, the term “responsibility to protect” appears once, and only in reference to the host state, with states grounding their votes in Chapter VII, not R2P. Second, there is no behavioral internalization: tested against the Finnemore–Sikkink norm life-cycle, a real norm would have states deciding by the rule rather than by cost–benefit, yet the P-5 immediately reverted to interest-driven inertia on Syria. Third, history shows it isn’t the first time, Southern Rhodesia (1965), Haiti (1994), and Somalia (1992) also saw force used for humanitarian ends without forming a stable norm. He quotes Bernard-Henri Lévy and Sarkozy, “the decision would not have occurred without the political will of one man”, to make the point that if an intervention turns on one leader’s personality or emotion, it cannot be a rule; even supporters saw it as exceptional. His method blends textual analysis, discourse analysis, historical comparison, and norm-theory testing, a constructivist-plus-historical-institutionalist evaluation. Is he convincing? The rebuttal is solid and persuasive but not a total refutation: he disproves that R2P was the main causal force, yet by his own admission it “possibly became one factor.” He shows the text–behavior disjuncture and uses mainstream internalization indicators, but his “negative-causation” method cannot prove R2P had no effect at all; it may still work as a background discursive frame. Supporters claimed Resolution 1973 was unique, the first time the Council authorized force for purely humanitarian purposes against a functioning state, giving it “precedential novelty.” Hehir answers that prior military actions induced similar optimism, so he picks two representative precedents: Southern Rhodesia (1965) and Haiti (1994). They prove three things: that Council-authorized intervention in response to domestic oppression long predated R2P, so Libya is neither the “first” nor unprecedented; that the Council acts on political feasibility, intervening only when P-5 interests temporarily align rather than when a norm automatically binds; and that the “precedential novelty” claim is hollow, Libya continues a pattern of selective consistency. Importantly, each such action was branded “exceptional” precisely to avoid creating precedents that would demand future consistency; the inconsistent use of Chapter VII angered many states and opened the Council to charges of hypocrisy. He cites Halderman on Rhodesia, better explained as “the result of a peculiar configuration of political forces and economic feasibility” than as consistently applicable law, and implies the same fits Libya. Are they good analogies? Yes, as evidence of a long-standing pattern of aberrant resolve amid inertia. On approach, Hehir is a constructivist, not a functionalist: a functionalist would judge the UNSC by institutional performance (does it efficiently solve cooperation problems, treating IOs as neutral tools); a constructivist judges R2P by its normative status, whether it has been internalized as a shared standard of legitimacy. Hehir asks the latter question and concludes R2P remains a political slogan rather than an embedded norm, a verdict driven by concern with norm diffusion and legitimacy, not institutional efficiency. This both converses with and pressures Beardsley & Schmidt: their data say crisis-mission factors usually drive UN involvement, while Hehir’s close reading of one high-profile case says even the apparent norm-triumph was really interest alignment. The tension is between an average statistical pattern and a single hard case, and both can be true at once.

Jo & Simmons (2016): Can the International Criminal Court Deter Atrocity?

International Organization · prosecutorial vs. social deterrence, tested on civilian killings

The ICC is the world’s first permanent and global criminal court, which helps set expectations about which wartime tactics fall beyond the pale. The research question is simply: can the ICC deter atrocity? It is extraordinarily hard from a research-design standpoint, and the literature is full of contradictory findings. Deterrence is by nature the study of non-events, how do you prove something didn’t happen because of deterrence? The classic identification problems all bite: selection and endogeneity (states that ratify or become ICC “targets” may differ systematically from those that don’t, they use coarsened exact matching and an instrumental-variable specification to probe this); temporal trends (violence may shift for reasons unrelated to the ICC, they include year/period dummies and find no particular trend in overall violence); reciprocity and the cycle of violence (government and rebel killing are mutually related, they control for the other side’s killings); and count-data quirks (killings are highly skewed and over-dispersed, with many country-years at zero, possibly a zero-inflation problem they don’t model explicitly, relying instead on the negative binomial’s flexibility). Data are also near-impossible to collect perfectly. They handle these well, distinguishing two mechanisms, positing conditional rather than all-or-nothing hypotheses, running many robustness checks, but the deep limit is that observational data can never fully solve causal identification: the ideal of randomly assigning which states join the ICC is impossible. Prosecutorial deterrence works via anticipated legalized, court-ordered punishment: actors omit crimes as the perceived odds or severity of legal sanction rise. Its mechanisms are Office-of-the-Prosecutor actions (investigations, warrants) that heighten the perceived likelihood of punishment, plus domestic legal reforms (complementarity) that raise enforcement capacity. Social deterrence results from extra-legal social costs of law violation: the ICC draws “bright lines” that mobilize domestic and foreign audiences to shame, isolate, and penalize violators. Its mechanisms run through ICC ratification interacting with the growth of human-rights organizations (mobilization) and with aid pressure (donor leverage). Which is more compelling? Their own estimates make prosecutorial signals, “ICC actions”, strong and robust for governments and even rebels, whereas social deterrence is conditional (stronger for ratifying states; stronger for secessionist, disciplined rebels). That pattern favors prosecutorial deterrence as the primary, general mechanism, with social deterrence as an amplifier where audiences and constraints exist. I’d add a worry about the social-deterrence chain, which is fragile: ICC investigation → public attention → government feels pressure → less violence presupposes a free press, a government that cares about reputation, and effective control of the military, preconditions that often fail in exactly the cases where the ICC intervenes (Sudan, Congo).

Table 2 (governments) is country-year: 2,264 observations, 101 countries, DV = the count of civilians intentionally killed by government forces. Table 3 (rebels) is rebel-group-year: 2,196 observations, 260 rebel groups, DV = civilians killed by rebel groups. Both use negative-binomial panel models with random effects plus year fixed effects; they prefer random effects to exploit both cross-country and over-time variation and to enable broader inference (RE assumes unobserved country traits are uncorrelated with ICC participation), while reporting fixed effects in the appendix as a robustness check (FE relaxes that assumption but uses only within-country change). The model fits over-dispersed count data and unit heterogeneity. Year fixed effects absorb common global shocks (end of Cold War, global democratization) but not unit-specific serial correlation, for which one could add a lagged DV or dynamic-panel methods; they also don’t model spatial spillovers, though ICC actions in one region could plausibly affect neighbors. Key takeaways, governments: ratification lowers killings (incidence-rate ratio ≈ .531); ICC actions (a three-year moving average of exams, investigations, warrants) reduce killings (IRR ≈ .570) beyond a simple post-ICC time effect; domestic statute reform (complementarity) reduces killings; and the social-deterrence interactions, ratification × HRO growth and ratification × aid pressure, are negative, ICC norms amplify social pressure. Rebels: ratification and domestic statutes have no general effect, but ICC actions still cut killings (IRR ≈ .830), even rebels “update” on investigations and warrants; and a weakly significant triple interaction (secessionist × disciplined × post-ICC) implies moderation among groups seeking legitimacy with command-and-control. Figure 1 standardizes this: with ’100 deaths’ as a no-effect baseline, ratification, ICC actions, and domestic statutes each move the expected count down for governments, while for rebels only ICC actions produce a clear reduction, derived from the IRRs in Tables 2–3. One clarification I’d make explicit in the caption: the ’100 deaths’ baseline is a didactic counterfactual multiplied by the IRR, not an empirical prediction for any specific case. Partly. The effect sizes aren’t huge, an IRR of .53 is a 47% reduction, but the baselines are very low (means of roughly 34 killings/year for governments, 83 for rebel groups), and because many cases are zero, statistical significance may be driven by a few extreme cases. The standout finding, rebels respond only to ICC actions, not to ratification, itself hints that prosecutorial deterrence outweighs social deterrence, yet prosecutorial deterrence faces the “can’t catch them” problem. Several alternatives remain: time-trend confounds (post-Cold War democratization, the general diffusion of human-rights norms beyond the ICC, more UN peacekeeping); reverse causation (violence falling first frees a state to ratify, rather than ratification cutting violence); and third factors (they control for aid and HRO growth but may omit regional-organization pressure, development level, electoral cycles). Measurement is shaky too: intentional-killing data come from media reports, which carry reporting bias (some countries draw more attention), classification difficulty (collateral damage vs deliberate killing), and large numbers of unreported cases. The deepest worry is external validity. The sample is mostly weak states that already voluntarily accepted ICC jurisdiction, with no constraint on the major non-members (US, China, Russia) who are precisely the key actors in the system, from a Chinese domestic vantage the ICC never seemed to have real bite. So the study really asks “among states that already accepted ICC jurisdiction, can its symbolic authority reduce marginal violence?”, not “can the ICC change the violent logic of great-power politics?”

Application · Theory in current events

Ovadek. 2025. “Hungary’s Exit from the International Criminal Court is a Sign of the Times.” The Conversation.

Ovadek reads Hungary’s withdrawal from the ICC as a symptom of a wider erosion of the rules-based order. The contemporary ICC already faces deep challenges: great-power non-membership and pushback (the US, Russia, and Israel signed but never ratified and later withdrew their signatures; China and India never signed), which leaves the Court without key cooperation and exposes it to political retaliation; an international environment less friendly to legal and judicial solutions, where even states that once “feigned” support now openly defy international law; an enforcement gap, since the Court depends on states to arrest and assist (South Africa’s failure to detain al-Bashir, Hungary’s treatment of Netanyahu); persistent charges of politicization and selective compliance; and ordinary resource and time constraints plus jurisdictional and sovereignty frictions (complementarity, consent requirements, the politics of aggression jurisdiction). Ovadek notes Europe was historically the ICC’s strongest backer, partly a Nuremberg legacy of faith in criminal accountability, partly the 1990s–2000s liberal project in which legalization and judicialization (WTO, ICC) were pillars of a rules-based order that aspiring EU members were rewarded for endorsing.

On Hungary specifically, Ovadek ties the exit to Hungary’s “kleptocratic authoritarianism” and its role as a “Trojan horse” for authoritarian interests inside the EU. The Netanyahu visit and the ICC warrant supplied the trigger, but the deeper driver is regime type and Orbán’s illiberal foreign-policy orientation. In my view this is a pattern, not a one-off: Ovadek situates it within Russia’s aggression, the US retreat from judicial mechanisms (WTO Appellate Body paralysis, threats against the ICC), and wavering signals from other European leaders, indicative of a broader trend rather than an isolated aberration. Does Barnett & Finnemore’s framework apply to the ICC? Yes. They see IOs as bureaucracies using rational-legal, moral, and expert authority, and the ICC fits: it classifies acts (war crimes, crimes against humanity, genocide), fixes meanings (the elements of crimes), and diffuses norms of accountability. It exercises power primarily through legal-bureaucratic authority and expertise, opening situations, issuing warrants, setting prosecutorial priorities, signaling standards that socialize states and armed actors. And it has changed over time: early optimism about legalization gave the Court strong symbolic and normative leverage, while today political pushback and selective cooperation constrain it, in B&F’s terms, classic IO pathologies (rigidity, insulation, proceduralism) now interact with state resistance to blunt its delegated authority.

The mirror-image question, should the US join the ICC?, pulls both course frameworks together. There is a credible case for at least deeper cooperation and removal of punitive measures: membership would bolster accountability, reduce double-standard charges, let the US shape rules from within the Assembly of States Parties (budget, elections, oversight), strengthen deterrence, and support allies who rely on the Court; the main worry, exposure of personnel, is partly answered by complementarity (which protects states with credible domestic prosecutions) and implementing legislation. Through Koremenos’s lens, US entry would expand membership legitimacy and ease enforcement; substantive scope stays fixed by the Rome Statute, though US participation could reshape priorities; centralization wouldn’t change formally, but US resources and secondments to the OTP could raise functional capacity; the US would gain real control in the Assembly of States Parties (classic principal-agent voice over agents); and although formal reservations are barred, amendment opt-ins (the aggression amendments), cooperation MOUs, and interpretive understandings would function as flexibility devices the US would push for. Through Jo & Simmons’s lens, US accession could strengthen both deterrence channels, prosecutorial deterrence via greater resources, cooperation, and credibility of investigations, and social deterrence by aligning US rhetoric with practice and amplifying the audience costs that mobilize against violators, though the marginal effect on the toughest cases (the great powers themselves) would remain limited, the same external-validity ceiling that constrains their original findings.


← Back to the genealogy tree  ·  ← Prev: Week 9  ·  Next: Week 11, Power Transition →