1 · Concept overview
Philosophy of science asks what science is, what makes its claims warranted, and what its terms refer to. The framing under test here is the one most readers arrive with without having chosen it: that there is a demarcation criterion separating science from non-science, that Popper supplied it, and that it works. This page is about why that is wrong, what replaced it, and what happens when a court has to apply it anyway.
Established Two positions the reader almost certainly holds, both standard, both rejected by the specialist literature, and both worth naming in the first screen. The first is that Popper settled demarcation with falsifiability. He did not: the criterion has known counterexamples in both directions, it requires a conventional decision at the base-statement level that Popper himself conceded, and forty years of argument since 1983 have produced no successor with agreed necessary and sufficient conditions. The second is that there is a scientific method — the one taught in school. No philosopher of science defends the five-step textbook procedure; the reference literature calls it a “legend”, describes it as “a fixed four or five step procedure starting from observations…progressing over formulation of a hypothesis…designing and conducting experiments…analyzing the results, and ending with drawing a conclusion”, traces it to Dewey's inquiry model and to Pearson, and says it misrepresents actual practice.
Frontier And the entry's own framing inverts the expectation the reader brings. Its position is that “the lack of a singular method across all disciplines of science is a reason for its reliability as a practice, not a mark against it”. That is a thesis rather than a measurement, and it is the shape of the whole subject: serious, well argued, and untested.
Frontier What the page lands is a two-part finding. The best-defended demarcation criterion classifies thoroughly refuted pseudoscience as science — and the more thoroughly refuted, the better it scores — while its own foundation is a communal convention its author described as a jury verdict. And a criterion the field has largely abandoned is quoted approvingly in a United States Supreme Court opinion that governs which expert testimony a jury hears. Both are true at once, and an honest treatment holds them together rather than resolving them.
Established A flagging note, because this subject is where the site's labels do the most work. On this page established means what a named author actually wrote, what a named court actually held, and what a named survey actually measured. A philosophical position with serious defenders and no empirical test is frontier or speculative however well argued and however widely held — and where a position is held by a majority of professional philosophers, that is an established fact about the profession and a frontier thesis about the world.
2 · Current scientific position
Established Popper's proposal is narrower and more careful than its reputation, and stating it accurately is the precondition for criticising it. It is a criterion of demarcation, not of meaning — Popper explicitly separated himself from the positivists here, holding that metaphysics is not nonsense, merely not science. It concerns logical form: a theoretical sentence is falsifiable if and only if it logically contradicts some empirical sentence, so a theory must prohibit something and scientific statements must be “capable of conflicting with possible, or conceivable observations”. And he condemned conventionalist stratagems — auxiliary hypotheses added solely to accommodate a failed prediction — which was his charge against Marxism and psychoanalysis, of which he wrote that “there was no conceivable human behaviour which could contradict them”.
Established The concession the reputation omits is Popper's own, and it is at the foundation. Basic statements — the singular observation reports that would do the falsifying — cannot themselves be justified by experience. Their acceptance is a matter of intersubjective decision by the research community, and Popper's own analogy is trial by jury: a verdict is a conventional agreement, not a correspondence with truth. The criterion built to replace communal agreement with logic rests, at its base, on communal agreement.
Established The failure is two-directional and this is the objection to lead with. Falsifiability admits pseudoscience: most pseudosciences are not unfalsifiable, they are falsified. Astrology, Popper's own showpiece, has been tested and thoroughly refuted (Culver and Ianna, 1988). On a purely syntactic criterion a refuted theory scores as more scientific, not less. And it excludes real science: theories not in the required prohibitive logical form, and any theory at a stage where its auxiliaries are unsettled, fall outside. Hansson (2006) makes both objections and the reference literature reports them as standard.
Established The reason falsification does not work as an algorithm is structural rather than a defect of Popper's formulation. Duhem, in 1914: “a physicist can never subject an isolated hypothesis to experimental test, but only a whole group of hypotheses”. A failed prediction indicts the conjunction, not any member of it, and Duhem's target was the crucial experiment — there are none, because the fault may lie in the auxiliaries. Established Duhem restricted the claim to theoretical physics, and that restriction is routinely dropped when the thesis is quoted; it is restored here. Quine generalised it in 1951 via the attack on the analytic–synthetic distinction: “the total field is so underdetermined by its boundary conditions, experience, that there is much latitude of choice as to what statements to reevaluate”. The reference literature further distinguishes holist underdetermination — which belief in a connected system to revise when evidence goes wrong — from contrastive underdetermination, whether evidence can favour one theory over empirically equivalent rivals, and insists these are “fundamentally distinct lines of thinking”.
Frontier The counter-argument is serious and the page carries it rather than letting underdetermination stand unopposed. Laudan and Leplin (1991) argue that empirical equivalence is not a permanent property: what counts as observable changes with instrumentation, auxiliary assumptions change, and equivalence judgements are therefore defeasible and time-relative. Philosophers have wrongly treated logically possible alternatives as epistemically equivalent ones. Kitcher's version is the practitioner's: “give us a rival explanation, and we'll consider whether it is sufficiently serious.” Whether that defeats underdetermination is unresolved.
Established Logical positivism did not fall to an external refutation. It was dismantled from inside, and the documentation is unusually complete. Ayer's 1936 formulation — a synthetic sentence is meaningful if it implies or adds to the observational content of some other sentence — admits any sentence whatsoever, since conjoining an arbitrary sentence with a conditional from that sentence to an observation yields a conjunction implying the observation. Hempel reviewed the crisis twice in a year, judging the criterion “lively and promising” in 1950 and reversing to unpromising by 1951 — a friendly party changing his verdict inside twelve months is the cleanest single datum on the criterion's condition. Carnap shifted testing from whole sentences to basic expressions in 1956; Kaplan produced counterexamples, published 1975; Creath showed in 1976 that they were repairable by modest revision. Quine removed the analyticity the whole apparatus rested on. The criterion was also self-undermining, not being itself verifiable. Political catastrophe and European emigration fragmented the movement's institutions, and Passmore declared it dead in 1967 while conceding a lasting legacy.
Established What replaced it is the correction most readers need, and it is neither “nothing” nor Kuhn. The reference account is that philosophy of science absorbed naturalism and engagement with actual practice; that probability theory developed in both frequentist and Bayesian directions, “the latter now dominant”; and that explication and conceptual engineering survived as tools. The institutional continuity is direct — nearly all US philosophers of science trace academic lineage to Reichenbach — so the movement's death “masks enduring influence”. The historiography of what came after belongs to Scientific Revolutions.
Established Laudan's 1983 argument is the pivot of the modern field, and its sharpest move is not the history. The structure: a history of failure — Aristotle, “to have science one must have apodictic certainty”; the seventeenth century dropping causal explanation and keeping certainty, with “Newton regarded his non-causal account as 'scientifical' because of the (avowed) certainty of its conclusions”; the nineteenth abandoning certainty for method, where “there was no agreement about what the scientific method was”; then verificationism and falsificationism. Then the claim that no proposed criterion is either necessary or sufficient. Then the syntactic trap: under falsificationism “flat Earthers, biblical creationists, proponents of laetrile or orgone boxes…all turn out to be scientific” provided they say something refutable. And the conflation that makes it fatal: “Scientific status, on their analysis, is not a matter of evidential support or belief-worthiness, for all sorts of ill-founded claims are testable and thus scientific.” A theory that fails its tests becomes more scientific — the exact inverse of the point of the exercise.
Established His recommendation is stronger than it is usually reported to be, and it is worth quoting exactly. “If we would stand up and be counted on the side of reason, we ought to drop terms like 'pseudo-science' and 'unscientific' from our vocabulary; they are just hollow phrases which do only emotive work.” The replacement question is not What makes a belief scientific? but What makes a belief well founded (or heuristically fertile)? Frontier Whether he is right is the live question of the field, and it is not settled by his paper's influence: it is simultaneously the most cited argument on the topic and the thing everyone since has been arguing with.
Frontier The responses are multi-criterial rather than replacement criteria, which is itself the field's answer. Pigliucci (2013) holds that the problem survives Laudan's attack. Hansson has pursued cluster accounts under the constraint that criteria be “applicable across disciplines with highly different methodologies”. Established Pennock (2011), in Synthese, is the most useful here: he argues that Laudan's demise thesis assumes an over-strict conception of what demarcation requires, and defends a “ballpark” criterion — methodological naturalism as a ground rule — that functions both in scientific bodies' explicit statements and in practice, and which he argues Laudan himself tacitly shares. Pennock testified in Kitzmiller, so he is an involved party writing about his own case, and that is marked rather than buried.
Frontier Realism is the other unfinished business, and readers meet its arguments as though they were settled. The no-miracles argument, Putnam: realism “is the only philosophy that doesn't make the success of science a miracle”. The pessimistic meta-induction, Laudan (1981): a historical list of empirically successful theories later rejected with their central terms judged non-referring — phlogiston, caloric — so by induction current theories will go the same way. Structural realism, Worrall (1989): be a realist about structure rather than about the natures of unobservables, because later theories preserve structural relations while revising the entities. Entity realism, Hacking (1982, 1983): believe in unobservables you can causally manipulate to intervene elsewhere, with warrant scaling with manipulative precision. And the base-rate objection, Lewis (2001) and Magnus and Callender (2004): the pessimistic induction commits the base-rate fallacy, because inferring from past failures requires a base rate of non-referring theories that is not available. Frontier Every one of these is a thesis with named defenders and no test.
Established There is one genuine measurement in this part of the subject, and it measures philosophers. The 2020 PhilPapers Survey, 1,785 respondents overall: on scientific realism against anti-realism, 72.35% accept or lean toward realism against 15.04% anti-realism (N = 1,689); on laws of nature, 54.31% non-Humean against 31.27% Humean (N = 1,554). That is citable and solid. It settles nothing about theories, and treating a professional majority as evidence for a thesis is the specific error this page is built to prevent.
Established And then demarcation acquired legal consequences. In Daubert v. Merrell Dow Pharmaceuticals, 509 U.S. 579 (1993), the Supreme Court displaced the Frye general-acceptance test with Rule 702 and gave trial judges a gatekeeping role. From the opinion, on testing: “Scientific methodology today is based on generating hypotheses and testing them to see if they can be falsified; indeed, this methodology is what distinguishes science from other fields of human inquiry.” The Court cited Popper. That is the single most consequential sentence any philosopher of science has ever had adopted by a court, and it was written the decade after the field's most influential paper argued the criterion does not work. Established The opinion is more careful than its reputation on the other factors: publication “is but one element of peer review” and “is not a sine qua non of admissibility; it does not necessarily correlate with reliability”; the court “ordinarily should consider the known or potential rate of error…and the existence and maintenance of standards controlling the technique's operation”; general acceptance is demoted to a factor, with a technique attracting “only minimal support within the community” properly viewed with scepticism; and the inquiry envisioned by Rule 702 is “a flexible one”.
Frontier The other two cases bracket it, and the more interesting one is the attack from the winning side. In McLean v. Arkansas Board of Education (E.D. Ark. 1982) Judge William Overton defined science by five essential characteristics — guided by natural law; explanatory by reference to natural law; testable against the empirical world; conclusions tentative and not necessarily the final word; falsifiable — with Michael Ruse testifying for the plaintiffs and grounding the test. Laudan then attacked an opinion he had every political reason to welcome, on the ground that creation science is falsifiable and has been falsified, so the correct verdict is bad science rather than non-science, and winning on a bad criterion stores up trouble. Frontier His commentary was paywalled and could not be obtained for this pass; the characterisation is second-hand and should be checked against the original before being quoted. Established In Kitzmiller v. Dover Area School District (M.D. Pa. 2005) Judge John E. Jones III held that intelligent design “is not science” on three grounds quoted in the opinion: that it “violates the centuries-old ground rules of science by invoking and permitting supernatural causation”; that irreducible complexity “employs the same flawed and illogical contrived dualism that doomed creation science”; and that “ID's negative attacks on evolution have been refuted by the scientific community”.
Frontier The sequel matters because it widened the doctrine. General Electric v. Joiner (1997) let judges exclude on analytical-gap grounds; Kumho Tire v. Carmichael (1999) extended gatekeeping to all expert testimony rather than only the scientific; Rule 702 was amended in 2000 to codify the trilogy. Frontier A 2002 RAND study is reported to have found a significant rise in exclusion of scientific expert testimony, with summary judgment against plaintiffs succeeding roughly twice as often — taken from an encyclopedia summary rather than from the RAND report, which was not obtained, and flagged accordingly. Frontier The standing tension is the page's terminal position rather than a paradox to be dissolved: Laudan says the demarcation vocabulary is hollow and emotive; American courts use it to decide what children are taught and which expert a jury hears. Both are true.
Established There is one more body of genuine measurement available to this page, and it measures the machinery of inference rather than the philosophers arguing about it. When the same data are handed to many independent teams, the analysis is not a pass-through. Twenty-nine teams and sixty-one analysts given one football dataset and one question returned odds ratios from 0.89 to 2.93, with twenty teams finding a significant effect and nine not; seventy teams given one neuroimaging dataset and nine pre-specified hypotheses produced no two identical analysis pipelines, and on several hypotheses a substantial minority reached the opposite conclusion from the majority. Established When the design is fixed before the data instead, the yield changes: comparing standard psychology articles with registered reports, positive results fall from about 96% to about 44%. And when replication is attempted at scale, 97 of 100 original studies had reported significant results against 36% of their replications, with mean effect sizes roughly halving. Frontier They bear directly on the demarcation dispute in a way no argument in this section does, because they show that the properties a criterion would have to detect — testability, error rate, the discipline of prediction — vary enormously within fields everyone agrees are sciences. The interesting variance is not between science and pseudoscience. It is between two laboratories in the same department.
3 · Frontier questions
Frontier Is any demarcation criterion possible? Laudan's answer is no and the vocabulary should be dropped; Hansson's and Pigliucci's is that multi-criterial or cluster accounts survive; Pennock's is that a ballpark criterion — methodological naturalism as a ground rule — is what scientific bodies actually state and what courts actually apply. Forty years of argument, no agreed necessary and sufficient conditions, and no test that could distinguish the positions. This is the field's central open question and it is open in the strong sense: nobody has proposed an observation that would settle it.
Frontier Does underdetermination bind in practice or only in logic? Duhem restricted his thesis to theoretical physics; Quine generalised it; Laudan and Leplin argue that empirical equivalence is defeasible and time-relative, so the space of genuine rivals shrinks as instrumentation improves. That is an empirical-sounding claim about the history of theory choice and nobody has tested it. A study asking how often an apparently empirically equivalent rival was later separated by new instrumentation would be evidence rather than argument, and would be the first of its kind.
Frontier Is the pessimistic meta-induction a base-rate fallacy? Lewis and Magnus and Callender say yes: inferring from a list of past failures requires knowing the proportion of successful theories that were non-referring, and that base rate is not available. The interesting feature of this objection is that it is a request for a number that could in principle be constructed — a systematically sampled inventory of empirically successful theories with their referential fate, rather than a curated list of the famous ones. Speculative No such inventory exists and building it would require settling in advance what counts as success, which is the disputed part.
Frontier What do practising scientists actually use? The reference literature is clear that published work does not describe the process — “scientific publications do not in general reflect the process by which the reported scientific results were produced”, with Medawar's Is the scientific paper a fraud? as the classic statement and the observation that “scientists' experimental practices are messy and often do not follow any recognizable pattern”. Method is treated as “detailed and context specific problem-solving procedures”. Frontier This pass could not find a systematic survey asking practising scientists what they take to demarcate science — the empirical question the framing most needs answered. The nature-of-science education literature and the Mertonian norms surveys are adjacent and are not this. A documented absence is a finding, and it is recorded as a typed requirement in section 13 rather than smoothed over. The norms evidence itself is filed in Scientific Institutions Through History and not restated here.
Frontier The strong programme, stated in its own terms before it is answered. Bloor's four tenets: causality, social factors are causally relevant to belief formation; impartiality, true and false beliefs alike require causal explanation; symmetry, the same explanatory resources apply whatever the outcome; reflexivity, the sociology of knowledge applies to itself. Barnes, Bloor, Collins, MacKenzie, Pickering and Shapin studied how large-scale social phenomena settle scientific controversies. Established The reference literature states flatly that the programme “is often interpreted as claiming all science is merely social construction with no epistemic warrant” and that “actually, it proposes that social factors causally influence even justified beliefs — not that justification disappears”. That is the misreading this site's readers will arrive with.
Speculative And the misrepresentation runs the other way too, which is the half that gets left out. The programme is also under-read as harmless. Latour and Woolgar's laboratory ethnography was taken, in later work, to show “that philosophical analyses of rationality, of evidence, of truth and knowledge, were irrelevant to understanding scientific knowledge”. Shapin and Schaffer's Leviathan and the Air-Pump (1985) traced how Boyle's experimental practice was socially legitimated against Hobbes rather than adjudicating who had the better argument. That is a stronger claim than Bloor's four tenets, it is separate from them, and answering the weak version does not answer it.
Frontier The serious responses are not dismissals, and three of them are worth the page's space. Longino's contextual empiricism holds that the cognitive processes of science are themselves social and that objectivity is a property of communities meeting four norms — venues for critical interaction; uptake of criticism evidenced by belief change over time; public accessibility of standards; and “tempered equality of intellectual authority”, where standing can be lost by failure to engage. Solomon's social empiricism: “A community is rational when the theories it accepts are those that have unique empirical successes” — and, strikingly, “individual irrationality can contribute to community rationality” through an appropriate distribution of biases. Kitcher's division of cognitive labour (1993) holds that a community should maintain a balance of orthodox and maverick strategies, and that “scientific progress can tolerate and indeed benefits from a certain amount of 'impure' motivation”.
Speculative Kitcher's well-ordered science (2001) is the most ambitious and the most exposed. Research priorities should track what a suitably constituted representative body would decide under ideal deliberation with expert input — criticised for its counterfactual form, for assuming public funding dominance, for national boundedness, and for separating research direction from research conduct. Frontier The field then splits three ways on what social influence means: the isolationist reply (Laudan, Brown, Goldman, Haack) that social factors are distorting, so “when evidential considerations have not trumped non-evidential considerations, we have an instance of bad science”; the reconciliationist reply (Giere's satisficing, Kitcher's External Standard) that social context is compatible with preserved epistemic norms; and the integrationist worry that social influence may be constitutive rather than distorting, while constructivists underspecify how communities distinguish knowledge from opinion — “a genuine philosophical problem left unresolved”.
Frontier Does adversarial collaboration actually resolve anything, and what does it produce when it does not? The procedure — rival proponents agreeing in advance on a design, running it jointly, and publishing together — was championed by Kahneman, who reported the discouraging part himself: participants rarely concede, and the usual outcome is that each side finds the result compatible with its prior position. That is a finding about the method and not a reason to abandon it, because the valuable output is not agreement. It is a public record of which predictions each side made before the data existed, which is exactly what the ordinary literature does not preserve. Frontier Whether such records accumulate into anything — whether a proponent who is wrong three times running loses standing, in Longino’s sense of an authority that can be lost by failure to engage — has never been measured, and the number of published adversarial collaborations is small enough that it could be: count them, score the predictions, and see whether anybody’s position moved. Established A weaker version of the same instrument is already known to work: prediction markets and structured surveys forecast which published findings would replicate at about 70% accuracy, well before the replications ran. A field can therefore see its own failures coming and does not, at present, do anything with that information.
4 · Technological bottlenecks
Frontier The binding constraint in this subject is that its central claims have no test, and this is a structural feature rather than a temporary state. Whether falsifiability demarcates, whether Scheffler's referentialism defeats incommensurability, whether the no-miracles argument survives the pessimistic induction, whether Longino's four norms constitute objectivity — none of these is a claim about which an observation could be decisive, and none of their defenders claims otherwise. That is why almost nothing on this page carries the established flag except as a report of what someone wrote.
Frontier The one place a measurement is both possible and missing is the operative criteria of practising scientists. The philosophical literature has produced a dozen candidate demarcation criteria in eighty years, and nobody has asked the population that would have to apply them what they use. The instrument is not hard — a forced-choice and free-response survey across disciplines, benchmarked against the published criteria — and the sampling frame exists. Its absence means the field's central question has no data on the one population whose answer would matter.
Established The second bottleneck is a documentation asymmetry that shapes what this page could say. The Daubert opinion was fetched in full and every quotation from it here is verbatim from the opinion text. McLean and Kitzmiller were not: the five Overton characteristics and Jones's three grounds are taken from encyclopedia summaries quoting the opinions, and are named as such. Laudan's 1982 commentary attacking McLean is paywalled and was not obtained. The result is that the case with the most consequence is the one this page can quote and the two creationism cases are the ones it must attribute at one remove — the opposite of what a page about demarcation in court would choose.
Frontier The third is that the field's operative reference layer is a single encyclopedia. Almost every attribution on this page — Popper's logical form, Duhem's page 187, Quine's pages 42–43, Ayer's collapse, Hempel's reversal, Putnam's page 73, Bloor's tenets, Longino's four norms — runs through peer-reviewed entries rather than through the primary texts, because the primary texts are monographs in copyright. The entries are excellent and are cited as entries. The dependency is still real and is stated rather than hidden.
Frontier And the fourth is legal rather than philosophical. Daubert's error-rate factor asks for “the known or potential rate of error” and for “standards controlling the technique's operation”. This pass located no published error rate for the operative task in several forensic disciplines routinely admitted under Rule 702. A criterion the opinion states plainly is unmet in practice by evidence the same opinion governs, and that gap is typed as an institutional requirement in section 13.
5 · Research dependencies
Established Nothing on this map produces a result this brief waits on, and no typed depends-on edge is claimed. What it waits on is one measurement nobody has made — a survey of what working scientists take to demarcate science — and one institutional condition nobody has supplied: published error rates and operating standards for the forensic disciplines admitted under the rule Daubert governs. Both are recorded as typed requirements below.
Established The seams with the three adjacent briefs are stated here so neither side restates the other. Scientific Revolutions owns Kuhn, incommensurability, Lakatos's research programmes and the historiography of change. Popper appears in both, and the seam is: this page asks whether falsifiability separates science from non-science; that page asks whether change by refutation matches the historical record, and hands the answer to Lakatos. Laudan splits the same way — the 1983 demarcation argument and the 1981 pessimistic induction are here, his place in the post-Kuhnian revolt is there. Scientific Institutions Through History owns Merton's norms and the survey evidence on whether scientists follow them; this page points there rather than restating it. Scientific Governance Models owns peer review as an allocation rule, registered reports, replication policy and integrity machinery; this page touches peer review only where Daubert quotes it as a demarcation factor.
Established Sources this page needs and did not obtain, named rather than paraphrased: Laudan's 1982 Science at the Bar commentary, paywalled, its argument carried second-hand; the McLean and Kitzmiller opinions themselves, quoted here through encyclopedia summaries; the 2002 RAND study on post-Daubert exclusion rates, taken from a tertiary summary and flagged frontier where its figures appear; and any systematic survey of practising scientists' demarcation criteria, which does not appear to exist in a form this pass could find.
6 · Required experiments
Frontier The highest-value missing study in this subject is a survey, and it is cheap. Ask a stratified sample of practising researchers across disciplines what would make them classify a claim as outside science; present the published criteria — falsifiability, methodological naturalism, testability, tentativeness, community acceptance — and measure endorsement, disagreement between fields, and the gap between endorsement and application on worked cases. The result would be a fact about the world rather than a position in a debate, and it would be the first one this field has had on its central question.
Frontier Second: measure what Daubert did. The doctrine has thirty-three years of trial records, a codifying amendment in 2000 and a reported RAND finding this page will not rely on. A systematic study of admissibility outcomes by field, with the exclusion rationale coded against the four illustrative factors, would establish which factor is actually load-bearing — and whether the Popper-citing sentence does any work at all or is decorative relative to general acceptance, which the opinion demoted but which judges may still be applying.
Frontier Third: build the base rate the pessimistic induction needs. Sample empirically successful theories systematically rather than by reputation, fix in advance what counts as success and what counts as a central term, and record referential fate. Speculative The design is contestable at exactly the point that matters and that is the interesting part: an attempt that failed publicly on its sampling frame would still be more informative than the current position, which is two curated lists pointed at each other.
Frontier Fourth: operationalise Longino's four norms. Venues for criticism, uptake evidenced by belief change over time, publicly accessible standards, and tempered equality of intellectual authority are all, in principle, observable properties of a research community. A field scored on the four and compared with its replication record would test the claim that objectivity is a community property — the single most testable proposal in the social-epistemology literature, and untested.
Speculative Fifth: test Solomon's counterintuitive claim directly. If individual irrationality can contribute to community rationality through an appropriate distribution of biases, then a community with a wider spread of prior commitments should converge faster or more accurately on questions with known answers. That is simulable now and partially observable in fields with settled controversies, and it is the one claim here whose failure would be informative to working scientists rather than only to philosophers.
7 · Engineering requirements
Frontier The only place this subject has engineering requirements rather than arguments is the courtroom, and they are stated in the opinion. Daubert asks a judge to consider whether a technique can be and has been tested, whether it has been subjected to peer review and publication, “the known or potential rate of error”, “the existence and maintenance of standards controlling the technique's operation”, and the degree of acceptance in the relevant community — under an inquiry that is “a flexible one”. Three of those five are properties a discipline can be engineered to have; the other two are properties of a community.
Frontier The error-rate factor is where the engineering gap is largest. A stated error rate requires a defined task, a ground-truth set, a blinded protocol and a published validation study. This pass located no published validation of that shape for several forensic disciplines routinely admitted under Rule 702. Building them is proficiency-testing infrastructure — laboratory accreditation, blind trials embedded in casework, published performance distributions — and it is a funding and standards problem rather than a research problem.
Frontier The second engineering fact is that peer review is doing work in this doctrine that the opinion itself declines to give it. Daubert says publication “is but one element of peer review” and “is not a sine qua non of admissibility; it does not necessarily correlate with reliability”. The Court was more cautious about publication as a reliability signal in 1993 than much of the practice built on its opinion has been since. How that machinery performs is Scientific Governance Models's subject and this page hands off to it.
Frontier And the third is that a survey instrument is engineering too. The missing measurement of operative demarcation criteria fails on instrument design rather than on cost: asking scientists what demarcates science invites them to recite the school method, which the reference literature calls a legend. The instrument would have to be behavioural — worked cases with forced classification and stated reasons — rather than declarative, and the norms-survey literature shows exactly why: scientists endorse norms they do not report following.
Frontier The method this corpus can actually build, stated as a protocol rather than as an aspiration. Six components, each answering a specific failure recorded above. One, rival hypotheses with named proponents. A contested question is written as a pair of claims, each attached to a person or a school willing to be identified with it, because the alternative — a literature review that reports a disagreement in the passive voice — is what allows a dispute to persist for forty years without a test. Two, a preregistered crux. Before any evidence is gathered, each side records the observation that would move it and by how much. Three, a shared dataset and a declared analysis plan. The many-analysts results are the reason: a shared dataset without a shared plan produces a spread of answers wide enough to contain both parties’ positions, so the plan must be fixed first, and where it cannot be, the honest form is to run the analysis several ways and publish the spread as the result.
Frontier Four, an independent replication arm. Whoever ran the first analysis does not run the second, and the replication is commissioned before the first result is known, so the decision to check does not depend on the answer. Five, scored forecasts. Both sides, and anyone else who wishes, record a probability on the stated crux with a resolution date; the probabilities are published with names attached and scored when the question resolves. This is the component that makes the rest cumulative, because it converts a rhetorical position into a track record, and a track record is the only mechanism in this entire literature that could make being wrong cost anything. Six, visible unresolved disagreement. The output is not a consensus paragraph. It is a joint statement of what both sides now agree the evidence shows, what they still dispute, and — the part normally omitted — why the evidence did not settle it: the auxiliaries were unsettled, the measurement was not sensitive enough, or the parties were using a term differently. Established That last category is Duhem’s and Quine’s thesis being used constructively rather than defensively: the conjunction was indicted, and saying which member of it is now in question is a result.
Frontier The failure modes are known in advance and belong in the protocol rather than in the post-mortem: attrition, a crux whose resolution date outlasts anyone’s attention, and theatre between proponents who were never far apart. All three show up as a forecast record that never resolves, which is what the scoring registry is for. Handwave No protocol of this shape has been run at the scale of a research programme, the registered-report and many-analysts evidence supports three of the six components and not the other three, and the corpus should expect at least one of them to prove unworkable as written.
8 · Adjacent technologies
Within this map: Scientific Revolutions, which owns Kuhn, incommensurability and the historiography of change, and shares Popper and Laudan with this page along the seam stated in section 5; Scientific Institutions Through History, which owns Merton's norms and the only direct measurements of whether scientists follow them; Innovation History, which owns the measurement-validity problems in the quantitative study of science; and Scientific Governance Models, which owns peer review, replication policy and integrity machinery as they operate now.
Also within it: Future Legal Systems, where the gatekeeping question this page ends on is a live design problem; Scientific Advisory Institutions, which inherits the demarcation problem in the form of deciding whose expertise counts; Artificial Scientists, where “method is domain-specific problem-solving procedures rather than an algorithm” becomes an engineering constraint; and Future Education Systems, which is where the school scientific method is actually taught.
Outside it: evidence law and the Rule 702 trilogy; forensic science standards and proficiency testing; the sociology of scientific knowledge; and science education research, which has its own literature on the nature of science and is adjacent to the missing survey without being it.
9 · Institutional requirements
Frontier This is the one subject on the map where a philosophical dispute has a named institution applying it under oath. American trial judges have been gatekeepers of expert evidence since 1993, applying illustrative factors drawn from a philosophical literature that had, ten years earlier, produced its most influential paper arguing the central criterion does not work. The institutional requirement is not a better criterion. It is a decision procedure a non-specialist can apply consistently under time pressure, which is a different object and which nobody in philosophy of science is building.
Frontier The second institutional fact is that the criterion's failure has a direction, and the direction has a constituency. Laudan's objection to McLean is that winning on a bad criterion stores up trouble — a philosopher on the winning side publishing against the judgment. That is unusual behaviour and it is the reason the commentary is cited here despite the page not having obtained its text. What is being defended is the idea that the reason a court gives has consequences beyond the case, and the record since supports it: the Rule 702 trilogy widened gatekeeping to all expert testimony, not only the scientific.
Established The interested parties on this page are marked where it matters. Pennock, whose 2011 paper is the best-sourced response to Laudan here, testified for the plaintiffs in Kitzmiller and is writing about his own case. The PhilPapers survey is a survey of a profession conducted by members of that profession, and is cited as a measurement of philosophers' opinions and nothing else. The encyclopedia entries are peer-reviewed reference works maintained by the field they describe, cited as entries rather than as the philosophers they report.
Frontier And a structural point about the field's own institutions. Logical empiricism's collapse is usually told as a story about arguments. The reference account gives arguments and adds political catastrophe: European emigration fragmented the movement's institutions, and the succeeding generation of US philosophers of science traces academic lineage overwhelmingly to one emigre. A field whose dominant position changed partly because its institutions were destroyed and rebuilt around particular people is a case for its own subject matter, and it is not usually treated as one.
Frontier The reason this method is rare is institutional and entirely mundane. An adversarial collaboration costs two laboratories a joint project, takes longer than either would take alone, and its most valuable outcome — a crux that failed to resolve, with the reason stated — is the least publishable object in science. Every incentive in the system rewards the side that publishes first and separately. Frontier What an institute publishing briefs can do unilaterally is narrower and still worth doing: it can host the register. A standing record of contested claims, each with a named crux, a resolution date and dated probabilities from whoever wants to be scored, requires no laboratory, no funding agency and nobody’s permission — and it is the piece the literature is missing, because journals publish findings and nobody keeps score. Established This corpus already runs a crude version of the instrument without scoring it. Every flag on every page is a claim about evidential status, dated and attributable, and a page that says a question will be settled by a particular observation has made a forecast in everything but name. Scoring those retrospectively — how often did the frontier claims resolve toward established, and in which direction — is the cheapest available test of whether this house’s judgement is calibrated, and until it is run the flags are an editorial convention rather than a measured one.
10 · Ethical & societal considerations
Frontier Laudan's recommendation has a cost he acknowledged and his critics press. Dropping “pseudo-science” and “unscientific” as hollow emotive phrases removes the vocabulary in which public bodies refuse to teach creationism, license a therapy, or admit an expert. His answer is that the replacement question — what makes a belief well founded — does the same work better, because it asks about evidence rather than about category membership. Frontier Whether an institution can operate on that question is untested, and the evidence from the courts is that they reach for category membership anyway.
Frontier The syntactic trap has a direct ethical consequence that is easy to miss. If scientific status is a matter of logical form rather than of evidential standing, then a claim becomes more scientific by being refuted — and a movement can acquire the label by making a bold falsifiable claim and losing. That is not a philosophical curiosity; it is a strategy, and the demarcation vocabulary is what makes it pay.
Frontier The strong programme is misrepresented in both directions and both misrepresentations do damage. Read as denying epistemic warrant, it becomes a licence for treating scientific claims as merely political, which the reference literature says flatly is not what it proposes. Read as harmless, it obscures that later laboratory-studies work did claim philosophical analyses of rationality and evidence are irrelevant to understanding scientific knowledge — a strong claim that deserves an answer rather than a shrug. Stating each position in its own terms before answering it is the only defensible way to handle this, and it is why both versions appear above.
Frontier And the professional-majority error deserves its own line because it is committed constantly and in good faith. That 72.35% of surveyed philosophers accept or lean toward scientific realism is a solid measurement of a profession. It is not evidence that theories are approximately true. The same reasoning that makes a survey of philosophers uninformative about unobservables makes a survey of any expert community uninformative about the object of its expertise, and it is the same move as citing consensus in place of the evidence the consensus rests on.
11 · Civilizational implications
Frontier The terminal position is that the framing survives only in a much weaker form than it is usually held. Not there is a demarcation criterion, but there are defeasible, multi-criterial, domain-relative judgements that courts and communities in fact make and must go on making. That is Pennock's ballpark position and Hansson's cluster position, and neither claims what the framing claims. Nothing here licenses the conclusion that the distinction between science and non-science is empty; what the record shows is that it has never been formalised and that its formal versions fail in both directions.
Frontier The practical consequence runs through institutions rather than through epistemology. Every body that has to decide what counts as science — a court, a curriculum authority, a regulator, a funder, a platform moderating health claims — is applying a criterion the field cannot supply. They apply one anyway, and the evidence from the one case with a documented doctrine is that they end up with a flexible multi-factor test and considerable judicial discretion. The measured record on discretion versus rule in scientific institutions is Scientific Governance Models's, and it favours rules.
Speculative The long-run risk is that the demarcation question gets automated before it is answered. Classifying claims as scientific or not is exactly the kind of task large-scale systems are now asked to perform at volume, on health information, on research integrity screening, on grant triage. A criterion nobody can state will be approximated by whatever the training data encodes, which will be the school method and the general-acceptance test the Supreme Court demoted in 1993. That is a hypothesis about where the operative criterion will come from, not a prediction about capability.
Frontier And the durable finding is about method rather than about demarcation. The reference literature's position — that the absence of a single method across disciplines is a reason for science's reliability rather than a mark against it — inverts the popular understanding completely. If it is right, the thing to protect at civilisational scale is the plurality of domain-specific problem-solving procedures rather than a canonical procedure, and every reform that standardises method across fields is trading away the property that made the enterprise work. It is a thesis with no test, held by most of the specialists, and it is the most consequential unfalsified claim on this page.
12 · Timelines
These horizons track what could be measured or decided, not what philosophers will conclude:
- 10 yr: Frontier The demarcation debate does not resolve, because nothing in it is the kind of thing a decade settles. What could change inside the window is empirical: a survey of practising scientists' operative criteria is a doable study, and forensic proficiency-testing infrastructure is under active construction in several jurisdictions. Frontier Expect the courts to move faster than the philosophy, as they have since 1993, and expect the operative criterion in practice to remain general acceptance under a different name.
- 25 yr: Speculative The plausible development is that demarcation migrates from philosophy to institutional design: not what is science but what decision procedure should a non-specialist gatekeeper apply, which is answerable and is currently nobody's research programme. Speculative Social epistemology is the likeliest source of the first testable proposal, because Longino's norms and Solomon's bias-distribution claim are the only positions here that make observable predictions about communities.
- 50 yr: Speculative If automated classification of claims becomes the operative gatekeeper at scale, the effective demarcation criterion becomes whatever those systems encode, and the philosophical question becomes an auditing question. Handwave That is an extrapolation from a deployment pattern to an epistemic outcome, with no evidence that the deployment will take the form assumed.
- 100 / 250+ yr: Handwave Beyond useful forecasting. The one datum at this horizon is that the criteria used to characterise science have turned over roughly every century since Aristotle — certainty, causal explanation, method, verifiability, falsifiability — which is Laudan's argument and is a periodisation rather than a rate. Handwave Five turnovers in two thousand years is a story, not a base rate.
Frontier One addition, because it is the only item here with a date attached. A scored crux register produces nothing in its first year and its first real output arrives when the earliest resolution dates fall due — which means the method is falsifiable on a timetable, unlike almost everything else on this page. Handwave If after a cycle of resolutions no proponent has changed position and no forecast record has altered anyone’s standing, the correct conclusion is that the instrument measures without disciplining, and that is worth knowing too.
13 · Technology tree & dependencies
- Depends on Nothing on this map. No brief here produces a result this one waits on, and no typed depends-on edge is claimed. The constraints that bind are a measurement nobody has made and a standards regime nobody has built, both typed below.
- Requires (not on this map) A stratified survey of practising researchers' operative criteria for classifying a claim as outside science — a specific instrument, not the brief's title restated. Present worked cases, force a classification, record the stated reason, and benchmark against the published criteria: falsifiability, methodological naturalism, testability, tentativeness, community acceptance. Measure endorsement, between-field disagreement, and the gap between what respondents endorse and what they apply, which the norms-survey literature says will be large. This pass could not find such a study in any form, and it is the empirical question the whole demarcation debate most needs answered; it is typed as a missing result rather than an institutional gap because the sampling frame and the funding already exist and nobody has run it. And published error rates and operating standards for the forensic disciplines admitted under Rule 702: the Daubert opinion asks a judge to weigh “the known or potential rate of error” and “the existence and maintenance of standards controlling the technique's operation”, and this pass located no published error rate for the operative task in several disciplines routinely admitted. That is proficiency-testing infrastructure — defined tasks, ground-truth sets, blinded protocols embedded in casework, published performance distributions — and it is a standards and funding decision rather than a discovery, which is why it is typed as institutional. And a scored crux register for contested claims: a standing record in which a disputed question is written as rival hypotheses with named proponents, a crux each side agrees would move it, a resolution date, and dated probabilities that are scored when the question resolves. It is typed here because nothing in it is a discovery — forecasting tournaments and replication markets have shown the scoring works at around 70% accuracy — and because the missing ingredient is a body willing to keep the record and publish the scores, which no journal, funder or learned society currently does.
- Enables Scientific Advisory Institutions, Future Legal Systems and Scientific Governance Models all contain bodies that must decide whose claims count as scientific, and all of them currently do so without a criterion this field can supply. No typed enabling edge is claimed, because supplying the criterion is exactly what this page reports has not been achieved — an edge would assert a deliverable that does not exist.
- Adjacent Evidence law and the Rule 702 trilogy; forensic science standards and proficiency testing; the sociology of scientific knowledge; science education research on the nature of science; and within this map Scientific Revolutions, Scientific Institutions Through History and Innovation History.
14 · Common misconceptions & speculative claims
Established “Popper settled demarcation.” He proposed a criterion of demarcation and never of meaning; it fails in both directions, admitting refuted pseudoscience and excluding theories whose auxiliaries are unsettled; and he conceded that the basic statements it falsifies with cannot be justified by experience, their acceptance being an intersubjective decision he compared to a jury verdict. Frontier Forty years since Laudan's paper have produced no successor with agreed necessary and sufficient conditions, and the serious surviving positions are all multi-criterial.
Handwave “The scientific method is the five-step procedure.” This is the clearest case on the site of a claim defended by nobody in the relevant field and believed almost universally outside it. The reference literature calls it a legend, quotes the four-or-five-step formulation to reject it, traces it to Dewey and Pearson rather than to scientific practice, and holds that publications do not in general reflect the process that produced the results — Medawar's Is the scientific paper a fraud? being the classic statement. Frontier The specialist position is stronger than the correction: the absence of a single method is offered as a reason for science's reliability rather than an embarrassment.
Established “Falsifiability can be applied to a hypothesis on its own.” Duhem denies exactly this: a physicist can never subject an isolated hypothesis to test, only a whole group, so a failed prediction indicts the conjunction and there are no crucial experiments. Established Duhem restricted the thesis to theoretical physics and the restriction is usually dropped in quotation; Quine's generalisation is a separate and stronger claim resting on the attack on analyticity. Frontier Laudan and Leplin's reply — that empirical equivalence is defeasible and time-relative — is the best counter and is not a refutation.
Established “Logical positivism was refuted from outside, by Kuhn or by the sociologists.” It was dismantled from inside. Ayer's own formulation admitted every sentence whatsoever; Hempel reversed his published assessment within twelve months; Carnap's successive repairs bought counterexamples that Creath then showed were repairable again; Quine removed the analyticity the apparatus rested on; and the criterion was not itself verifiable. Frontier And what replaced it was naturalism, engagement with practice and Bayesian probability — with nearly all US philosophers of science tracing lineage to Reichenbach. The movement's death masks enduring influence.
Frontier “Laudan's demise thesis is the field's settled verdict.” It is the field's most influential paper on the topic and the thing everyone since has been arguing with, which are different properties. Pigliucci holds the problem survives; Hansson pursues cluster accounts; Pennock argues Laudan assumes an over-strict conception of what demarcation requires and defends methodological naturalism as a ballpark ground rule he says Laudan tacitly shares. Speculative None of these positions has a test, and the disagreement between them is not the kind that evidence resolves.
Established “Courts have solved what philosophers could not.” Daubert itself says the Rule 702 inquiry is “a flexible one” and that publication “is not a sine qua non of admissibility”; McLean's five-part test drew a published attack from a philosopher on the winning side; and the doctrine has since widened to all expert testimony rather than converging on a criterion. Frontier What courts have is a procedure and a decider, which is what an institution needs and what a criterion is not.
Established “The strong programme claims science has no epistemic warrant.” The reference literature states that this is the standard misreading, and that the programme proposes that social factors causally influence even justified beliefs rather than that justification disappears. Speculative Nor is the programme therefore harmless, and this is the half usually left out: later laboratory-studies work was taken to show that philosophical analyses of rationality, evidence, truth and knowledge are irrelevant to understanding scientific knowledge. That is a separate and much stronger claim, and answering the weak version does not touch it.
Established “72% of philosophers are realists, so realism is probably right.” The figure is real — 72.35% accept or lean toward scientific realism against 15.04% anti-realism, N = 1,689, in a survey with 1,785 respondents overall — and it is evidence about philosophers. Frontier The arguments it aggregates are the no-miracles argument against the pessimistic meta-induction, and the second is charged with the base-rate fallacy by Lewis and by Magnus and Callender. A majority holding a position no one can test is a fact about a profession.
Frontier “Social constructionism about science has been refuted, and Kuhn refuted it.” Kuhn rejected the constructivist appropriation of his own book, which is a fact about Kuhn and not an assessment of the programme; that assessment is a separate matter with its own literature, set out in section 3. Frontier The serious replies are three and none of them is a refutation: isolationism, that social factors are distorting and their dominance marks bad science; reconciliationism, that social context is compatible with preserved epistemic norms; and integrationism, which grants that social influence may be constitutive while noting that constructivists underspecify how communities distinguish knowledge from opinion — described in the reference literature as a genuine philosophical problem left unresolved.
Frontier And the framing itself: “there is a demarcation criterion” is a claim about logic doing duty as a claim about evidence. The best-defended criterion is syntactic, and a syntactic criterion severs scientific status from evidential standing — which is why flat-earthers and orgone boxes pass it and why a thoroughly refuted astrology scores better than a young theory whose auxiliaries are unsettled. Established Beneath that sits Popper's jury verdict, and beneath that sits Duhem: falsification as an algorithm applied to a single hypothesis does not exist. The honest replacement is Laudan's question — what makes a belief well founded — and no institution on the record has yet been built to answer it.