Statistical Reporting in CausalPy#

This page explains the statistical concepts used in CausalPy’s reporting layer. The reporting functions automatically compute and present statistics appropriate to your model type.

Model Types and Statistical Approaches#

CausalPy supports two modeling frameworks, each with its own statistical paradigm:

Model Framework

Statistical Approach

Statistics Reported

PyMC models

Bayesian

Mean, Median, HDI, Tail Probabilities, ROPE

Scikit-learn models

Frequentist (OLS)

Mean, Confidence Intervals, p-values

Note

The reporting layer automatically detects which type of model you’re using and generates appropriate statistics. You don’t need to specify the statistical approach.

Experiment Support#

The effect_summary() method is available for the following experiment types:

Experiment Type

PyMC Models

Scikit-learn (OLS) Models

Difference-in-Differences

✅ Full support

✅ Full support

Regression Discontinuity

✅ Full support

✅ Full support

Regression Kink

✅ Full support

❌ Not implemented

Interrupted Time Series

✅ Full support

✅ Full support

Synthetic Control

✅ Full support

✅ Full support

PrePostNEGD

✅ Full support

❌ Not implemented

Instrumental Variable

❌ Not available

❌ Not available

Inverse Propensity Weighting

❌ Not available

❌ Not available

Note

For experiments marked with ❌, use the experiment’s .summary() method for results.


Bayesian Statistics (PyMC Models)#

When you use PyMC models, CausalPy performs Bayesian inference, yielding posterior distributions for causal effects. The reported statistics summarize these posterior distributions.

Point Estimates#

Mean

  • The average of the posterior distribution

  • Represents the expected value of the causal effect

  • When to use: Most commonly reported point estimate; balances all posterior information

  • Interpretation: “The average estimated effect is X”

Median

  • The middle value of the posterior distribution (50th percentile)

  • Divides the posterior probability mass in half

  • When to use: Preferred when the posterior is skewed; more robust to outliers

  • Interpretation: “There’s a 50% probability the effect is above/below X”

Important

For symmetric posteriors, mean and median are nearly identical. For skewed posteriors, they may differ substantially. Report both to give readers a complete picture.

Uncertainty Quantification#

HDI (Highest Density Interval)

  • A credible interval containing the specified percentage of posterior probability (default: 95%)

  • Reported as hdi_lower and hdi_upper in summary tables

  • The narrowest interval containing the specified probability mass

  • Interpretation: “We can be 95% certain the true effect lies between X and Y”

  • Key difference from CI: This is a probability statement about the parameter itself, not about the procedure

Note

Public effect_summary() methods take alpha, not hdi_prob: their HDI coverage is 1 - alpha. The default alpha=0.05 therefore reports a 95% HDI, while alpha=0.025 reports a 97.5% HDI. This is independent of the project-wide HDI_PROB = 0.94 setting used by other reporting and plotting APIs, so one experiment can legitimately show a 95% summary HDI and 94% plot bands.

Example interpretation:

mean: 2.5, 95% HDI: [1.2, 3.8]

“The estimated effect is 2.5 on average, and we can be 95% certain the true effect lies between 1.2 and 3.8.”

Posterior Tail Summaries#

direction selects the tail summary that effect_summary() reports; it does not infer a direction from the posterior mean and does not change the HDI+ROPE conclusion.

  • direction="increase" reports p_gt_0 = P(effect > 0).

  • direction="decrease" reports p_lt_0 = P(effect < 0).

  • direction="two-sided" reports p_two_sided = 2 × min(P(effect > 0), P(effect < 0)) as a two-sided tail probability.

The table also retains prob_of_effect = 1 - p_two_sided for compatibility. This is a tail-derived score, not a posterior probability that the effect is non-zero: with a continuous posterior, P(effect != 0) is generally one regardless of that score. Tail summaries are descriptive evidence, not a binary credible/not-credible verdict.

ROPE and Practical Significance#

min_effect is the smallest absolute effect that is practically meaningful. It must be finite and non-negative; 0 is valid and creates the point ROPE [0, 0].

The p_rope table column#

For compatibility, the optional p_rope column is a direction-sensitive strict exceedance probability:

  1. direction="increase": P(effect > min_effect).

  2. direction="decrease": P(effect < -min_effect).

  3. direction="two-sided": P(|effect| > min_effect).

It is not the probability mass inside the ROPE and is not the input to CausalPy’s prose conclusion.

The HDI+ROPE conclusion#

When you supply min_effect, CausalPy instead uses the closed, symmetric ROPE [-min_effect, min_effect] and compares the reported HDI [L, U] with it.

  • If U < -min_effect or L > min_effect, the HDI is entirely outside the ROPE: the effect is practically significant.

  • If -min_effect <= L and U <= min_effect, the HDI is entirely inside the ROPE: the effect is practically equivalent to zero.

  • Otherwise the HDI overlaps the ROPE: the result is inconclusive.

Because the ROPE is closed, an HDI that only touches -min_effect or min_effect is inconclusive unless it is wholly inside. The report gives the posterior mass below, inside, and above the ROPE; samples exactly on a boundary count as inside, so those three masses partition the finite posterior draws. For time-series summaries with cumulative=True, CausalPy reports the same decision and mass breakdown separately for the average and cumulative effects. Non-finite posterior draws are excluded consistently from the reported mean, median, HDI, tail summary, ROPE masses, and conclusion; CausalPy raises ValueError if no finite posterior draws remain.

Without min_effect, effect_summary() remains strictly descriptive: it reports the interval, the requested tail summary, and any requested relative effect, but gives no binary or practical-significance verdict.

Scope of the HDI+ROPE verdict#

The HDI+ROPE verdict accounts for only part of the relevant uncertainty. It is a within-model statement: it propagates posterior uncertainty about the effect parameter conditional on the model, its priors, and the observed data—and nothing else. Thus, a practically significant, practically equivalent, or inconclusive verdict describes practical significance only within those modelling assumptions.

It does not represent design or identification uncertainty: whether the counterfactual is credible, whether parallel trends hold, whether a synthetic-control donor reconstruction holds, or how much a comparable period or unit moves when no intervention happened at all.

Where a valid, identified placebo null can be constructed, it is the better reference for total uncertainty. PlaceboInTime re-fits against pre-treatment placebo periods and learns a posterior-predictive status-quo null; PlaceboInSpace re-fits each control unit as if treated and returns the resulting placebo-unit effect summaries for comparison. In each case, empirical variation in periods or units presumed untreated adds design-relevant uncertainty that a posterior HDI alone does not contain. When those estimates provide an identified placebo null, that null captures strictly more of the total uncertainty than a posterior HDI. Read HDI+ROPE as the model-internal significance statement, but prefer judging practical significance against an identified placebo null. Agreement between the two is reassuring; disagreement is a signal that the model understates uncertainty.

PlaceboInTime does not always provide a null to compare against. It abstains with an INCONCLUSIVE result when fewer than two folds are usable or when the usable folds do not identify a between-fold spread, so there is no placebo null in either case. This learned null’s spread is the same quantity the sensitivity checks already report and forthcoming power-analysis work will consume, so these uses must remain coherent rather than becoming different notions of “significant.”

summary = result.effect_summary(direction="increase", min_effect=1.0)
print(summary.table["p_rope"])  # P(effect > 1.0)
print(summary.text)  # HDI+ROPE conclusion and below/inside/above masses

Important

Choose min_effect from domain knowledge: it should reflect the smallest effect that would justify the intervention cost or be scientifically meaningful.


Frequentist Statistics (Scikit-learn / OLS Models)#

When you use scikit-learn models (OLS regression), CausalPy performs classical frequentist inference based on t-distributions.

Point Estimates#

Mean / Coefficient Estimate

  • The estimated causal effect from the regression model

  • For scalar effects (DiD, RD): the coefficient of interest

  • For time-series effects (ITS, SC): the average or cumulative impact

  • Interpretation: “The estimated effect is X”

Note

Unlike Bayesian estimates, frequentist point estimates don’t come with a probability distribution. Uncertainty is captured through confidence intervals and standard errors.

Uncertainty Quantification#

Confidence Intervals (CI)

  • Reported as ci_lower and ci_upper in summary tables

  • Computed using t-distribution critical values at the specified significance level (default: α = 0.05 for 95% CI)

  • Interpretation: “If we repeated this experiment many times, 95% of such intervals would contain the true effect”

  • Key difference from HDI: This is a statement about the procedure, not about the parameter

Standard Errors

  • Measure of uncertainty in the coefficient estimate

  • Used to construct confidence intervals and compute p-values

  • Derived from the residual variance and design matrix

  • Larger standard errors → wider confidence intervals → more uncertainty

Example interpretation:

mean: 2.5, 95% CI: [1.1, 3.9]

“The estimated effect is 2.5. If we repeated this study many times, 95% of such confidence intervals would contain the true effect.”

Important

Bayesian HDI vs Frequentist CI: While numerically similar, they have fundamentally different interpretations. The HDI makes a direct probability statement about the parameter (“95% probability the effect is in this range”), while the CI makes a statement about the procedure (“95% of such intervals would contain the true parameter”).

Hypothesis Testing#

p-values

  • The probability of observing data at least as extreme as what we observed, assuming the null hypothesis (no effect) is true

  • Reported as p_value in summary tables

  • Common threshold: p < 0.05 is often used as evidence against the null hypothesis

  • Interpretation: Lower p-values indicate stronger evidence against no effect

Correct interpretation:

  • p = 0.03: “If there were truly no effect, we’d observe data this extreme only 3% of the time”

  • NOT: “There’s a 97% probability of an effect” (this is a Bayesian interpretation)

Common pitfalls to avoid:

  1. ❌ “p = 0.06 means no effect” → The p-value doesn’t prove the null hypothesis

  2. ❌ “p < 0.05 means the effect is important” → Statistical significance ≠ practical significance

  3. ❌ “p = 0.01 is better than p = 0.04” → Both provide evidence against the null; the effect size matters more

  4. ❌ “p > 0.05 proves no effect” → Absence of evidence is not evidence of absence

Decision guidance:

  • p < 0.05: Conventional threshold for “statistical significance”

  • p < 0.01: Stronger evidence against the null

  • p > 0.05: Insufficient evidence to reject the null (but doesn’t prove no effect)

Note

Always report the actual p-value and effect size, not just whether p < 0.05. The magnitude and confidence interval of the effect are often more informative than the p-value alone.

t-statistics and degrees of freedom

  • t-statistic = coefficient / standard error

  • Measures how many standard errors the estimate is from zero

  • Degrees of freedom (df) = n - p, where n = sample size, p = number of parameters

  • Larger |t-statistics| and smaller p-values indicate stronger evidence


Choosing Between Approaches#

When to use Bayesian inference (PyMC models):#

  • ✅ You want direct probability statements about effects

  • ✅ You have prior information to incorporate

  • ✅ You need uncertainty quantification for complex hierarchical models

  • ✅ You want to test against meaningful effect sizes (ROPE)

  • ✅ Small to moderate sample sizes where uncertainty matters

When to use Frequentist inference (OLS models):#

  • ✅ You need computational speed (OLS is faster)

  • ✅ Your audience expects classical statistical inference

  • ✅ Large sample sizes where approaches converge

  • ✅ Simple linear models without hierarchy

  • ✅ You want to align with traditional econometric practice

Important

Both approaches are valid and will often lead to similar conclusions, especially with larger sample sizes. The choice often depends on your field’s conventions, computational constraints, and whether you value direct probabilistic interpretation (Bayesian) or long-run frequency guarantees (frequentist).


Summary Statistics by Effect Type#

Scalar Effects (DiD, RD, Regression Kink)#

For experiments with a single causal effect parameter:

Bayesian output:

  • One row with: mean, median, hdi_lower, hdi_upper

  • Tail probabilities: p_gt_0 (or p_lt_0, or p_two_sided + prob_of_effect)

  • Optional: p_rope (if min_effect specified)

Frequentist output:

  • One row with: mean, ci_lower, ci_upper, p_value

Time-Series Effects (ITS, Synthetic Control)#

For experiments with multiple post-treatment time points:

Two aggregation levels:

  1. Average effect: Mean effect across the post-treatment window

  2. Cumulative effect: Sum of effects across the post-treatment window

Additional statistics:

  • Relative effects: Percentage change relative to counterfactual

    • relative_mean: Effect size as percentage of counterfactual

    • relative_hdi_lower / relative_hdi_upper (Bayesian)

    • relative_ci_lower / relative_ci_upper (frequentist)


Usage Examples#

Understanding the Output#

The effect_summary() method returns an EffectSummary object with two attributes:

Numerical Summary (.table):

Returns a pandas DataFrame with all statistics:

summary = result.effect_summary()
print(summary.table)

Prose Summary (.text):

Returns a human-readable interpretation ready for reports:

print(summary.text)
# Output: "The average treatment effect was 2.50 (95% HDI [1.20, 3.80]). The posterior probability of an increase is 0.975."

Basic usage (default Bayesian):#

import causalpy as cp

# Fit experiment with PyMC model
result = cp.DifferenceInDifferences(...)

# Get effect summary with default settings
summary = result.effect_summary()
print(summary.text)  # Prose interpretation
print(summary.table)  # Numerical summary

With directional hypothesis:#

# Test for an increase
summary = result.effect_summary(direction="increase")  # Reports p_gt_0

# Test for a decrease
summary = result.effect_summary(direction="decrease")  # Reports p_lt_0

# Two-sided tail summary
summary = result.effect_summary(direction="two-sided")  # Reports p_two_sided

With a practical significance threshold:#

# Define effects smaller than 2.0 in magnitude as practically equivalent.
summary = result.effect_summary(
    direction="increase",
    min_effect=2.0,
)
print(summary.table)  # Includes the legacy direction-sensitive p_rope column.
print(summary.text)  # Reports an HDI+ROPE verdict and posterior ROPE masses.

For time-series experiments with custom window:#

# ITS or Synthetic Control result
summary = result.effect_summary(
    window=(10, 20),  # Only analyze time points 10-20
    cumulative=True,   # Include cumulative effects
    relative=True      # Include percentage changes
)

Further Reading#

For deeper understanding of these statistical concepts:

  • Bayesian inference: The PyMC documentation provides excellent tutorials on Bayesian statistics

  • Causal inference: See our :doc:causal_written_resources for recommended books

  • Statistical terms: Refer to the :doc:glossary for concise definitions

  • Practical application: Explore the example notebooks in our documentation showing these concepts in action

See also

  • :doc:glossary - Quick reference for statistical terms

  • :doc:causal_written_resources - Books and articles on causal inference

  • API documentation for the effect_summary() method