Statistical Models

Hierarchical Bayesian models for AI evaluation data

Models in hibayes are NumPyro probabilistic models that define the joint probability distribution over random variables and observed data. There are a number of built-in models you can select from.

What makes up a model?

Models take processed Features and define a probabilistic program using NumPyro primitives (numpyro.sample, numpyro.deterministic, etc.). They return nothing but instead define the generative process for your data.

from hibayes.model import Model, model
from hibayes.model.models import check_features

import numpyro
import numpyro.distributions as dist
import jax.numpy as jnp
import jax


@model
def my_custom_model(
    prior_mu_loc: float = 0.0,
    prior_mu_scale: float = 1.0,
) -> Model:
    """
    A simple hierarchical model for grouped binomial data.
    """
    def model(features):
        check_features(features, ["obs", "num_group", "group_index", "n_total"])

        overall_mean = numpyro.sample(
            "overall_mean",
            dist.Normal(prior_mu_loc, prior_mu_scale),
        )

        sigma_group = numpyro.sample(
            "sigma_group",
            dist.Exponential(rate=1.0),
        )

        z_group = numpyro.sample(
            "z_group", dist.Normal(0, 1).expand([features["num_group"]])
        )
        group_effects = overall_mean + sigma_group * z_group
        numpyro.deterministic("group_effects", group_effects)

        logit_p = group_effects[features["group_index"]]

        numpyro.sample(
            "obs",
            dist.Binomial(
                total_count=features["n_total"],
                probs=jax.nn.sigmoid(logit_p),
            ),
            obs=features["obs"],
        )

    return model
1
Any prior hyperparameters the user can configure through the config file.
2
The inner function receives processed features from your data pipeline.
3
Use check_features to validate that required features are present, providing helpful error messages.
4
Use numpyro.deterministic to track derived quantities for posterior analysis.
5
The observed data must be named "obs" - this is enforced by the @model decorator.

Built-in models

Hierarchical binomial models

These models are designed for aggregated binomial data (success counts out of totals) with hierarchical grouping structures.

simplified_group_binomial_exponential - A two-level hierarchical model with an exponential prior on group standard deviation. Allows more group heterogeneity than the default two-level model, making population inference especially sensitive to priors when groups are few.

two_level_group_binomial - The package default. Uses half-normal priors on scale components and a non-centered parameterization. Its default HalfNormal(0.1) group standard deviation strongly favors similar groups; this is a substantive pooling assumption.

three_level_group_binomial - Three-level hierarchical model (group → subgroup → subsubgroup) with half-normal priors. Use for deeply nested data structures.

three_level_group_binomial_exponential - Three-level hierarchical model with exponential priors on variance components. Alternative prior specification.

Choosing a grouped binomial estimand

In the two grouped binomial models, a group’s success probability is sigmoid(overall_mean + sigma_group * z_group). overall_mean is the mean log-odds, so sigmoid(overall_mean) is the median probability across the modeled population of groups, not its arithmetic mean probability.

Choose the quantity that answers the evaluation question:

Estimand Question and weighting
median_group_probability What is the probability for a group at the population’s median? sigmoid(overall_mean); useful for interpreting the existing latent site.
evaluated_group_mean_probability What is average performance across the evaluated groups? Equal weight per group with positive trial count.
evaluated_trial_mean_probability What is average performance under this evaluation’s allocation of trials? Weight group probabilities by their total n_total, combining repeated rows.
population_mean_probability What is expected performance for a randomly drawn exchangeable group? Integrate sigmoid(overall_mean + sigma_group * Z) over Z ~ Normal(0, 1) for each posterior draw.

The last quantity describes a population average. Its posterior interval captures uncertainty about that average, not variation in the rate of an individual new group or future binomial outcomes. Population group weights are implicit in the exchangeable-group model; this is not an item-weighted deployment estimate unless that assumption matches the application.

Few observed groups weakly constrain population location and heterogeneity. For example, all-zero results may be explained by negative observed-group effects while the population location remains close to its prior. Adding trials to one group improves its rate estimate, but does not add independent information about other groups. Group-level estimates and population estimates may therefore differ substantially without a sampler error. Compare reasonable priors on both location and group standard deviation, including boundary data and different numbers of groups. A pooled-binomial reference assumes a different sampling model and is not a universal target.

Use the binomial estimands communicator to report these probabilities with posterior uncertainty, including from saved fits. The generic summary_table continues to summarize the original latent sites. Usecase1 compares several models; the exponential model’s position in that config does not make it the package default.

Linear models

linear_group_binomial - Non-hierarchical binomial regression with configurable main effects and interactions. Treats categorical variables as predictors without nesting structure.

ordered_logistic_model - Ordered logistic regression for ordinal outcomes. Supports main effects, interactions, and configurable cutpoint priors.

Ordinal priors for scores at the ends of the rubric

The original ordinal priors remain the default: intercept ~ Normal(0, 1), first_cutpoint ~ Normal(-4, 0.2), and positive cutpoint spacings from a LogNormal(-0.5, 0.3) plus 0.3. They favor scores spread across the rubric. In a model with no effects, they give only about 1.75e-17 prior probability to P(score = 0) >= 0.99. This can substantially inflate estimated nonzero rates after an all-zero evaluation. This is strong prior influence, not a fixed probability floor: sufficiently many observations can overcome it.

When nearly all scores being at either end is plausible, use the opt-in boundary preset:

model:
  models:
    - name: ordered_logistic_model
      config:
        tag: boundary
        num_classes: 11
        prior_preset: boundary
check:
  checkers:
    - prior_predictive_plot: {interactive: false}
    - r_hat
    - divergences
    - ess_bulk
    - ess_tail
    - bfmi
    - posterior_predictive_plot: {interactive: false}

The preset uses an intercept scale of 5 and centers the expected cutpoint span on zero. With K categories, the first cutpoint mean is -(K - 2) * (exp(prior_cutpoint_diffs_loc + prior_cutpoint_diffs_scale**2 / 2) + min_cutpoint_spacing) / 2. For two categories it is zero. Effect priors, spacing priors, and the first cutpoint scale remain unchanged.

Explicit prior_intercept_scale and prior_first_cutpoint_loc values override the preset, including when supplied through YAML. Leaving them unspecified uses the selected preset; existing explicit configurations retain their meaning.

The ordered-logistic likelihood is evaluated as categorical log probabilities using stable log-space identities. This is the same statistical likelihood; it avoids cancellation when cumulative probabilities round to one near the rubric ends, including with float32 arithmetic. Sample-site names are retained.

This is a substantive prior choice, not an uninformative setting. It gives more weight to extreme scores and can affect small-sample conclusions. It does not address arbitrary concentration in an interior category, and its baseline calibration does not guarantee appropriate predictions after adding effects or changing the rubric. Check predicted category frequencies and compare plausible prior choices for the actual evaluation. A posterior interval excluding exactly zero is not, by itself, a calibration failure.

Regression tests cover prior predictions for 2, 5, and 11 categories, all-zero posterior behavior across sample sizes, and likelihood and gradient stability at extreme predictors. These finite checks do not establish general coverage or replace validation for the actual evaluation.

Configuring models

Models are configured in your hibayes.yaml under the model section:

model:
  models:
    - name: two_level_group_binomial
      config:
        fit:
          samples: 2000
          warmup: 1000
          chains: 4
        prior_mu_overall_loc: 0.0
        prior_mu_overall_scale: 1.0

Or specify multiple models to compare:

model:
  models:
    - name: simplified_group_binomial_exponential
      config:
        tag: "exponential_prior"
    - name: two_level_group_binomial
      config:
        tag: "halfnormal_prior"

Custom models

To use a custom model, create a Python file with your model definition and reference it in your config:

model:
  path: path/to/my_models.py
  models:
    - name: my_custom_model
      config:
        prior_mu_loc: -1.0

The @model decorator registers your function so it can be accessed by name in the config.

Built-in Models Reference

Hierarchical Binomial Models

For aggregated success/failure data with grouped structure.

Model Description Variance Priors
simplified_group_binomial_exponential Two-level hierarchy (overall → group). Population estimates can be sensitive to priors with few groups. Exponential on group standard deviation
two_level_group_binomial Default model for grouped data; favors strong pooling under its default group scale. HalfNormal on standard deviations
three_level_group_binomial Three-level hierarchy (overall → group → subgroup → subsubgroup). For deeply nested data like students within schools within districts. HalfNormal on all variance components
three_level_group_binomial_exponential Three-level hierarchy with simpler exponential priors. Alternative when HalfNormal priors are too informative. Exponential on all variance components

Regression Models

For modelling effects of categorical/continuous predictors.

Model Description Key Parameters
linear_group_binomial Non-hierarchical binomial regression. Treats groups as fixed effects rather than random. Supports main effects and interactions with optional effect coding. main_effects, interactions, effect_coding_for_main_effects
ordered_logistic_model Ordered logistic regression for ordinal outcomes (e.g., 0-10 ratings). Models cumulative probabilities via ordered cutpoints. main_effects, interactions, num_classes, effect_coding_for_main_effects