← Reference · Nestor G Pestelos Jr

Philosophy

Steelmanning

Reference entry · last updated August 25, 2026

Steelmanning is the practice of constructing the strongest possible version of an argument, typically one held by an opponent, before evaluating or responding to it.[1] It is named as the deliberate inverse of the straw man, a fallacy in which an arguer substitutes a weaker version of an opponent's position for the position actually held.[2] This entry covers the term's origin, its debt to the older philosophical idea of the principle of charity, the manual and AI-assisted methods used to do it, and the fallacy it produces when a position is strengthened beyond what its holder would recognize or endorse.

Etymology

The word combines steel with the suffix -man, formed as the deliberate inverse of straw man: a figure built of straw is weak and easy to knock down, one built of steel is not. The earliest documented use in something close to its current sense appears in a footnote to a 2011 essay on disagreement by Luke Muehlhauser, published on the community blog LessWrong: "Sometimes the term 'steel man' is used to refer to a position's or argument's improved form … a 'steel man' is an improvement of someone's position or argument that is harder to defeat than their originally stated position or argument."[1] The footnote's phrasing, "sometimes the term is used," suggests the word was already circulating before the essay rather than being introduced there for the first time.

No earlier documented use has been established; some secondary sources attribute coinage to other individuals, but none point to a primary source that predates or corroborates that claim, so it is not repeated here.

Relationship to the principle of charity

Steelmanning descends from an older idea in the philosophy of language: the principle of charity, the interpretive norm that a speaker's statements should be understood as more likely rational and true than a hostile or overly literal reading would suggest. The idea is sometimes traced to a 1959 paper by Neil L. Wilson on reference, and it plays a role in W. V. O. Quine's account of radical translation, where assuming a speaker is broadly rational becomes a precondition for translating them at all.[3] Donald Davidson gave the principle its most influential statement: for Davidson, charity is not a matter of politeness toward a speaker but a condition for interpretation to succeed in the first place, since without assuming a speaker is largely rational and largely right, no coherent meaning can be assigned to their words at all.[4][5] Steelmanning applies the same interpretive move specifically to argument evaluation: charity toward what a position could mean, extended to charity toward the strongest case that position could make.

The steel man and the straw man

In informal logic, the straw man is a well-documented fallacy: an arguer distorts, oversimplifies, or substitutes a weaker version of an opponent's position, refutes that substitute, and treats the refutation as if it applied to the position actually held.[2] Scott Aikin and John Casey's influential treatment identifies three related distortions along the same axis: the straw man proper (misrepresenting the position), the weak man (accurately representing one of an opponent's actual positions, but a weak or unrepresentative one, and treating it as the position), and the hollow man (attacking a position no identifiable person actually holds).[2] Steelmanning is defined by moving in the opposite direction along all three axes at once: toward the most accurate, the strongest, and the most genuinely held version of the position under discussion.

Method

Manual protocol

A common form of the practice runs in five steps: listen to the position without formulating a rebuttal first; ask what the strongest evidence for it is and under what conditions its holder would consider it wrong; summarize the position back in its strongest form; check that summary against the position's actual holder, adjusting until they agree it is fair, or until the exact point of disagreement is isolated; and only then respond to that strengthened version, not an earlier or weaker one.

No single named source documents this exact five-step sequence; it is a synthesis of the practice as commonly described, not a citation to one text.

The ideological Turing test

Economist Bryan Caplan proposed a specific test of steelmanning ability by analogy with the Turing test for machine intelligence: a person passes an "ideological Turing test" for a position they disagree with if their stated version of that position is indistinguishable, to a neutral judge, from a version stated by someone who actually holds it.[6] Caplan introduced the test in response to a claim that people on one side of a political divide understand the other side's arguments better than the reverse is true, arguing that the ability to state an opposing view as clearly and persuasively as its own proponents is itself "a genuine symptom of objectivity and wisdom."[6]

LLM-assisted protocol

A language model can be used to generate the opposing case directly: build the strongest version of one argument first, then, in a separate session or a different model given only the argument itself and not the reasoning that produced it, prompt for the strongest possible opposite case. Keeping the second step in a fresh context matters in practice, because a model (or person) that helped construct the original argument tends to defend it even under an explicit instruction to attack it, while a reader encountering only the finished argument has no such investment.

No external source found for this specific fresh-context requirement as a named technique; it is a practical extension of the general method above, not a documented finding.

Applications

Adversarial review, a pattern in AI agent systems where an agent's output is checked by a separate process instructed to find fault with it rather than confirm it, is a structured, repeatable version of the same underlying move: engaging the strongest available case against a claim instead of the first available one. Formal debate between AI systems is a further step in the same direction. Irving, Christiano, and Amodei propose training two AI agents to argue opposing sides of a question in front of a human or AI judge, on the premise that finding a flaw in an opponent's argument is an easier problem to solve than constructing a correct answer unassisted, and that a debate format therefore surfaces a stronger case on each side than either side would produce alone.[7]

Limitations

The iron man fallacy. Aikin and Casey identify a failure mode on the other side of steelmanning: reconstructing a position as stronger than any version its actual holder has stated or would endorse, then treating a response to that invented, stronger version as if it settled the debate. They term this the iron man fallacy. It misrepresents the opponent's position just as the straw man does, in the opposite direction, and can make a debater's eventual concession look more generous, or an opponent's actual position look weaker, than either is.[8]

"Steelmanning your own argument" is a different thing. Steelmanning targets an opposing position; strengthening one's own existing claim, or explaining away evidence against it, is not the same activity even when it borrows the name. Motivated reasoning research finds that a reasoner's motivation to reach a particular conclusion changes which evidence and which reasoning strategies get used, while the reasoner still experiences the process as appropriately rigorous throughout.[9] A search for reasons a position is right, run only in that direction, is difficult to distinguish from motivated reasoning by introspection alone; genuine steelmanning is checked externally, against the actual opposing case, not against how convincing the exercise feels from the inside.

Scope. The practice presumes the argument under reconstruction is a genuine, good-faith position capable of a stronger form. Applied to a deliberately manipulative or bad-faith argument, strengthening it produces a more effective piece of manipulation.

No external source found for this specific scope limitation as a named finding; it follows from the definition of steelmanning as an act of good-faith interpretation, not from a documented study.

See also

References

  1. ^ Luke Muehlhauser ("lukeprog"), "Better Disagreement," LessWrong, October 24, 2011 — https://www.lesswrong.com/posts/FhH8m5n8qGSSHsAgG/better-disagreement
  2. ^ Scott F. Aikin & John Casey, "Straw Men, Weak Men, and Hollow Men," Argumentation 25(1), 87–105 (2011). DOI: 10.1007/s10503-010-9199-y
  3. ^ W. V. O. Quine, Word and Object (MIT Press, 1960)
  4. ^ Donald Davidson, "Radical Interpretation," Dialectica 27(3–4), 313–328 (1973). DOI: 10.1111/j.1746-8361.1973.tb00623.x. Free full text: marcellodibello.com (PDF)
  5. ^ "Donald Davidson," Stanford Encyclopedia of Philosophyhttps://plato.stanford.edu/entries/davidson/
  6. ^ Bryan Caplan, "The Ideological Turing Test," EconLog, June 20, 2011 — https://www.econlib.org/archives/2011/06/the_ideological.html
  7. ^ Geoffrey Irving, Paul Christiano, Dario Amodei, "AI Safety via Debate" (2018). arXiv: 1805.00899
  8. ^ Scott F. Aikin & John P. Casey, "Straw Men, Iron Men, and Argumentative Virtue," Topoi 35, 431–440 (2016). DOI: 10.1007/s11245-015-9308-5
  9. ^ Ziva Kunda, "The Case for Motivated Reasoning," Psychological Bulletin 108(3), 480–498 (1990). DOI: 10.1037/0033-2909.108.3.480. Free full text: fbaum.unc.edu (PDF)