1  What is QCA Robustness

Before interpreting a perturbation-based robustness diagnostic, a researcher using QCA should be able to complete two sentences:

The reference result was tested against changes in _____.

It was considered preserved when _____ remained unchanged or sufficiently similar.

The first blank identifies the analytical decision being varied. The second sheds light on the QCA object or relation used to judge preservation. Unless both are stated, calling a result “robust” tells us neither what the analysis withstood nor what, exactly, survived.

Exceptionally, the cluster and theory comparisons introduced later do not perturb one analytical decision along such a path. For those diagnostics, the corresponding questions are which sufficient relation is being evaluated within which groups, or which condition-set specifications are being compared.

1.1 Establish the Comparator

Most diagnostics in this book begin with a baseline reference analysis that establishes the population of cases, the calibrated outcome and condition sets, the inclusion and frequency cutoffs used to construct the truth table, the treatment of logical remainders, and the conservative, parsimonious, or intermediate solution formula selected for comparison. Although the examples do not produce multiple admissible models or intermediate branches, we always specify, for didactic purposes, which model position or branch will be followed when appraising robustness.

Oana and Schneider refer to the carefully constructed solution selected for substantive interpretation as the Initial Solution (Oana and Schneider 2024). This book usually speaks more broadly of the reference or baseline analysis because qcaERT often reconstructs more than the solution formula itself. Depending on the diagnostic, the comparison may also require the original calibration anchors, calibrated memberships, truth table, exclusions, parameters of fit, and case set.

Again, the theory-comparison diagnostic is the exception. theory.test() does not perturb one reference formula, but constructs a separate truth table and solution for each theoretically motivated condition set, then compares those analyses with one another. The common outcome, cases, truth-table cutoffs, solution-selection rules, and other shared settings make the comparison coherent, but none of the condition-set specifications is automatically entitled to serve as the reference.

The comparison must also involve an alternative worth considering. A calibration anchor can be moved to a mathematically feasible but conceptually indefensible value, just as a case can be removed even though the research question clearly places it inside the reference population. Either operation may change the solution formula, but neither automatically supplies meaningful evidence against the reference analysis. Skaaning’s early robustness exercises consequently focused on reasonable alternative calibration thresholds (Skaaning 2011), while Oana and Schneider distinguish empirically observed sensitivity ranges from the substantively plausible alternatives that should compose a harder robustness test (Oana and Schneider 2024).

Throwing every diagnostic at the wall and hoping the resulting heap looks reassuring is not a robustness strategy. The research design must tell us which decisions are genuinely uncertain, which alternatives remain defensible, and why the selected comparison bears on the claim being made.

1.2 Name the Part of the Analysis That Changes

Calibration, truth-table construction, case composition, and condition selection all disturb a different QCA object. They enter the workflow at different stages and can reach the solution formula through different mechanisms:

Analytical decision varied What changes first What may change afterward
Qualitative calibration anchor Cases’ memberships in one condition or in the outcome Membership in truth-table configurations, consistency of sufficiency, PRI, sufficiency (OUT) assignments, and the minimized solution formula
Frequency cutoff The empirical frequency required for treating a truth-table configuration as observed The distinction between observed configurations and logical remainders, followed by the minimized solution formula
Inclusion cutoff The consistency requirement for assigning qualifying truth-table configurations to the outcome OUT assignments and the configurations submitted to minimization as sufficient for the outcome
Case composition The cases contributing to configuration frequencies, consistency of sufficiency, PRI, and outcome assignments The truth table, logical remainders, and solution formula; if calibration is recomputed, the qualitative anchors and calibrated memberships may change as well
Condition set The number and meaning of the logically possible configurations The observed configurations, logical remainders, simplifying assumptions, prime implicants, and complete solution formula

As you can see, a formula change should be interpreted in light of the analytical operation that produced it. If raising the frequency cutoff turns a previously observed configuration into a logical remainder, the resulting formula change does not mean the same thing as one produced by removing a case. The former changes the empirical-representation requirement imposed on the existing configurational space; the latter rearranges the empirical information within it.

The distinction between one-at-a-time and joint changes is equally important. A calibration, frequency-cutoff, or inclusion-cutoff sensitivity path changes one analytical decision while holding the others fixed. An alternative-set appraisal can instead combine changes to several calibration anchors with different inclusion and frequency cutoffs. The sensitivity path identifies a local formula-preservation boundary for a single analytical decision, whereas the alternative-set appraisal examines what happens across multidimensional alternative specifications.

1.3 Name What Must Be Preserved

Once the analytical change has been identified, the researcher must decide which evidence will be used to evaluate it. Exact preservation of the solution formula, differences in its parameters of fit, similarity between the complete-solution memberships produced by two formulas, and changes in the cases’ set-theoretic positions are related questions, but they are not substitutes for one another (Wagemann and Schneider 2015; Oana and Schneider 2024).

1.3.1 The selected solution formula

Formula preservation focuses on whether the selected Boolean expression remains identical after the analytical change. The comparison must refer to the same solution type and, when relevant, the same selected model position and intermediate branch. The boundary diagnostics use this criterion to identify the last tested value that reproduces the reference formula and the first tested value that changes it or otherwise prevents the selected result from being compared.

“Formula preservation” is exact and easy to state, but it is actually rather tricky. Two different formulas can produce highly similar complete-solution memberships across the cases, while an identical formula can acquire different parameters of fit after the cases’ calibrated memberships or the composition of the case set changes.

1.3.2 Solution consistency, PRI, and coverage

When the same solution formula is reproduced, its relation with the calibrated outcome may nevertheless change because the cases or their calibrated memberships have as well. Solution consistency evaluates the degree to which membership in the complete solution is a subset of membership in the outcome. Solution PRI penalizes membership in the solution that is shared with both the outcome and its negation. And solution coverage evaluates how much membership in the outcome is covered by the union of the solution’s sufficient terms.

The meaning of a parameter comparison depends on the diagnostic. For instance, altset.test() compares solution consistency, PRI, and coverage only when the reference analysis and an alternative specification contain the same selected formula key—its fit comparison is therefore like for like. loo.test() and subsample.test(), by contrast, compare the parameters attached to the selected solution type and model position in the complete-data and reduced-data analyses whenever both results are available. Those parameters can be compared even when the corresponding formulas differ. In that situation, the reported difference describes how the parameter of fit changes across two different sufficient sets; it does not isolate a change in the fit of the reference formula itself.

Structured comparisons require a different interpretation. On the one hand, cluster.test() holds a selected solution formula or prime implicant fixed and recalculates its consistency and coverage within each cluster. On the other hand, theory.test() constructs a separate truth table and performs a separate minimization for each condition-set specification, so differences in solution consistency, PRI, or coverage refer to independently obtained solutions. The numerical columns may retain familiar names throughout the package, but the analytical comparison those numbers represent must still be made explicit.

1.4 Keep Research Design Attached to the Diagnostic

Thomann and Maggetti distinguish QCA as a formal data-analysis technique from QCA as a broader research approach (Thomann and Maggetti 2020). Calibration, truth-table construction, tests of necessity and sufficiency, and logical minimization belong to the technique; as concept formation, case and condition selection, scope conditions, the use of case knowledge, and the movement between theory and evidence belong to the broader research approach.

This distinction does not prohibit robustness appraisals involving cases or conditions. It tells us that such appraisals reach further into the research design than moving an inclusion cutoff within an otherwise fixed analysis. Deleting a case may change configuration frequencies, the available logical remainders and the parameters of fit; replacing a condition changes the theoretical specification and reconstructs the entire configurational space. Neither operation should be interpreted as a generic stress test detached from the reasons for selecting the original cases or conditions.

The evidentiary value of a diagnostic also depends on whether the study is case-oriented or condition-oriented, whether it emphasizes substantive theoretical interpretability or redundancy-free solutions, and whether its reasoning is primarily exploratory or deductive (Thomann and Maggetti 2020). For example, repeatedly changing case composition may be informative in a condition-oriented analysis whose larger case set contains genuine uncertainty about individual observations or delimitations. In a small-N, case-oriented study with carefully justified scope conditions, mechanically deleting every case does not acquire substantive meaning merely because the operation can be automated. The diagnostic must remain connected to the study’s case-selection rationale and intended inference.

1.5 Report the Denominator, Not Only the Rate

Every rate has a numerator and a denominator. While the numerator counts the comparable analyses that meet the preservation criterion, the denominator enumerates all analyses in which that criterion could actually be evaluated. A formula-preservation rate of 75%, for example, may mean that 60 of 80 comparable analyses preserve the baseline formula—not that 75 of all 100 attempted analyses do so. Analyses that fail or do not produce a comparable solution must be reported separately, and not counted as formula changes.

Transparency in QCA requires enough information for readers to reconstruct the analytical decisions that produced the truth table and solution, including calibration, case selection, truth-table cutoffs, the treatment of logical remainders, parameters of fit, and meaningful robustness tests (Wagemann and Schneider 2015). A robustness report must add the comparison design: which analytical decision changed, which choices remained fixed, which solution type, model position, or intermediate branch was monitored, and which criterion defined preservation.

The denominator is part of that account. Comparison availability tells us whether the intended formula or parameter comparison could be made at all. An attempted analysis may fail during calibration, truth-table construction, or minimization. One solution type may also remain available while another does not. For the affected solution type, a noncomparable result is neither formula preservation nor formula change and must not be counted as either one.

A formula-preservation rate should therefore report the number of comparable analyses as its denominator. A rate concerning solution consistency, solution PRI, or solution coverage must use the denominator established by the diagnostic’s fit-comparison rule. For example, altset.test() requires a common selected formula key, whereas loo.test() and subsample.test() require the relevant selected solutions and their parameters of fit to be available. These denominators need not be identical, and neither should be silently replaced by the total number of attempted analyses.

Unsuccessful and unavailable comparisons should not disappear from the report. Their frequency and causes may reveal that the tested alternatives often produce no sufficient truth-table configuration, no requested solution type, or another analytical failure. They should, however, remain separate from the numerator and denominator used to describe preservation among completed comparisons.

1.6 Follow the Source of Uncertainty Through the Book

The implementation chapters organize the workflow around the analytical decision or comparison that motivates the diagnostic:

  • Boundary diagnostics move calibration anchors, the frequency cutoff, or the inclusion cutoff along ordered paths. The alternative-set diagnostic then samples combinations of those choices within a specified search space.

  • Case-composition diagnostics reconstruct the QCA after deleting one case or drawing repeated subsamples, with calibration either held fixed or re-estimated when the research design calls for it.

  • Structured comparison diagnostics answer related but different questions. cluster.test() retains a selected sufficient expression and recalculates its consistency and coverage within substantively meaningful groups, while theory.test() gives each theoretically motivated condition set its own truth table and minimization before comparing the resulting formulas, parameters of fit, and complete-solution memberships.

None of these diagnostics can confer a universal seal of robustness. Together, they allow us to identify varied analytical decisions, preserved QCA objects or relations, points at which the reference result changed, and attempted comparisons that could not be completed.