<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article
  PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD with MathML3 v1.2 20190208//EN" "JATS-journalpublishing1-mathml3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="en">
<front>
<journal-meta><journal-id journal-id-type="publisher-id">METH</journal-id><journal-id journal-id-type="nlm-ta">Methodology</journal-id>
<journal-title-group>
<journal-title>Methodology</journal-title><abbrev-journal-title abbrev-type="pubmed">Methodology</abbrev-journal-title>
</journal-title-group>
<issn pub-type="ppub">1614-1881</issn>
<issn pub-type="epub">1614-2241</issn>
<publisher><publisher-name>PsychOpen</publisher-name></publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">meth.20361</article-id>
<article-id pub-id-type="doi">10.5964/meth.20361</article-id>
<article-categories>
<subj-group subj-group-type="heading"><subject>Original Article</subject></subj-group>

<subj-group subj-group-type="badge">
<subject>Data</subject>
<subject>Code</subject>
</subj-group>

</article-categories>
<title-group>
    <article-title>Nuances of Information Criteria for Bayesian Psychometric Models</article-title>
<alt-title alt-title-type="right-running">Information Criterion Nuances</alt-title>
    <alt-title specific-use="APA-reference-style" xml:lang="en">Nuances of information criteria for Bayesian psychometric models</alt-title>
</title-group>
    
    
<contrib-group content-type="authors">
    <contrib id="author-1" contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid" authenticated="false">https://orcid.org/0000-0001-7158-0653</contrib-id><name name-style="western"><surname>Merkle</surname><given-names>Edgar C.</given-names></name><xref ref-type="aff" rid="aff1"/><xref ref-type="corresp" rid="cor1">*</xref></contrib>
<contrib contrib-type="editor">
<name>
    <surname>Nájera Álvarez</surname>
    <given-names>Pablo</given-names>
</name>
<xref ref-type="aff" rid="aff2"/>
</contrib>
    <aff id="aff1"><institution>Department of Psychological Sciences, University of Missouri, Columbia, MO, USA</institution></aff>
    <aff id="aff2">Universidad Pontificia Comillas, Madrid, <country>Spain</country></aff>
</contrib-group>
<author-notes>
    <corresp id="cor1">219 McAlester Hall, Columbia, MO 65211, USA. <email xlink:href="merklee@missouri.edu">merklee@missouri.edu</email></corresp>
</author-notes>
    
<pub-date pub-type="epub"><day>31</day><month>08</month><year>2026</year></pub-date>
<pub-date pub-type="collection" publication-format="electronic"><year>2026</year></pub-date>
<volume>22</volume>
<issue>3</issue>
<fpage>195</fpage>
<lpage>224</lpage>
<history>
<date date-type="received">
<day>16</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>03</month>
<year>2026</year>
</date>
</history>
<permissions><copyright-year>2026</copyright-year><copyright-holder>Merkle</copyright-holder><license license-type="open-access" specific-use="CC BY 4.0" xlink:href="https://creativecommons.org/licenses/by/4.0/"><ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref><license-p>This is an open access article distributed under the terms of the Creative Commons Attribution 4.0 International License, CC BY 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p></license></permissions>
<abstract>
<p>It is common practice to compare Bayesian psychometric models via information criteria such as DIC and WAIC. Especially because these criteria can be automatically computed by MCMC software, it is easy to ignore the intricacies related to their computation. This often leads researchers to use noisy criteria that may lead to suboptimal analysis decisions. In this paper, we first review different forms of Bayesian information criteria that could be computed for psychometric models. We then consider best practices, highlighting computational pitfalls that can occur even when one is attempting to follow best practices. Finally, we provide recommendations for the metrics’ practical uses. The paper is intended to clarify conflicting recommendations from the literature and to raise awareness about ways that information criteria can behave unexpectedly.</p>
</abstract>
    <kwd-group kwd-group-type="author"><kwd>Bayesian information criteria, Bayesian SEM, cross-validation, DIC, MCMC, PSIS-LOO</kwd></kwd-group>

</article-meta>
</front>
<body>
    <sec sec-type="intro" id="intro"><title/>
<p>Bayesian statistical models are commonly compared via information criteria such as the Deviance Information Criterion (DIC; <xref ref-type="bibr" rid="bid__36">Spiegelhalter et al., 2002</xref>) and the Widely Applicable Information Criterion (WAIC; <xref ref-type="bibr" rid="bid__45">Watanabe, 2010</xref>). These metrics are especially popular because they are convenient: software will often automatically compute the metrics, and the popular decision rule of “select the model with the lowest value” is as simple as it gets.</p>
<p>Convenient and simple metrics do not always lead to optimal outcomes. This is especially true when we apply Bayesian information criteria to models with “random” parameters, including psychometric models with latent variables and mixed models with random effects. In these situations, it is possible to compute multiple values of the information criteria for the same model, with these multiple values potentially leading us to select different models. The difference lies in whether or not the random parameters (latent variables or random effects) are counted alongside other model parameters. While <xref ref-type="bibr" rid="bid__36">Spiegelhalter et al. (2002)</xref> mentioned this issue in their original DIC paper (referring to it as the model “focus”), the issue has often been overlooked in the time since.</p>
    <p>Software packages that supply Bayesian information criteria vary in what is computed, leading researchers to report different metrics without realizing it. The <italic>blavaan</italic> package (<xref ref-type="bibr" rid="bid__22">Merkle et al., 2021</xref>; <xref ref-type="bibr" rid="bid__24">Merkle &amp; Rosseel, 2018</xref>) computes information criteria using likelihoods that are marginal over random parameters. But for some complex models, the criteria require a large amount of post-estimation computation time. The Mplus software often computes DIC using marginal likelihoods, but it switches to conditional (on random parameters) for some complex models that would require long computation time (see <xref ref-type="bibr" rid="bid__1">Asparouhov &amp; Muthén, 2020</xref>). The Blimp software (<xref ref-type="bibr" rid="bid__14">Keller &amp; Enders, 2023</xref>) often computes DIC using conditional likelihoods, but it switches to marginal for some multilevel models. The <italic>brms</italic> package (<xref ref-type="bibr" rid="bid__5">Bürkner, 2017</xref>) computes information criteria using conditional likelihoods. The same is true of models programmed in BUGS, JAGS, Stan, or custom MCMC algorithms, because the model specification nearly always involves the random parameters. Further, the specific DIC computations differ across software because there exist multiple definitions of DIC in the literature (e.g., <xref ref-type="bibr" rid="bid__30">Plummer, 2008</xref>). It is not necessarily a problem that these multiple metrics exist, but it is a problem that researchers do not know the differences between them. It is often easy to tell whether or not the random parameters are involved in likelihoods by looking at each criterion’s “effective number of parameters” value: if this is close to the frequentist parameter count, then a marginal likelihood is probably being used. If this is close to the number of observations in the dataset, then it is more likely (but not a certainty!) that a conditional likelihood is being used.</p>
<p><xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref> considered these issues in factor analysis and simple item response models, recommending use of marginal likelihoods (marginal over latent variables) when computing information criteria. They also showed that use of conditional likelihoods (conditioning on latent variables) can lead to suboptimal theoretical properties and large Monte Carlo error. These results followed on related work by <xref ref-type="bibr" rid="bid__16">Li et al. (2016)</xref>, <xref ref-type="bibr" rid="bid__26">Millar (2018)</xref>, and <xref ref-type="bibr" rid="bid__42">Vehtari et al. (2016)</xref>. Recent work that is closer to psychometrics includes <xref ref-type="bibr" rid="bid__8">Du et al. (2023)</xref>, who considered Bayesian information criteria of multilevel models with missing data, <xref ref-type="bibr" rid="bid__10">Graves and Merkle (2022)</xref>, who considered how identification constraints influence Bayesian information criteria, and <xref ref-type="bibr" rid="bid__47">Zhang et al. (2019)</xref>, who considered various versions of DIC for multilevel IRT models.</p>
<p>The goal of this paper is to provide further clarification and discussion on Bayesian information criteria for psychometric models, and to examine related problems in more complex psychometric models. With regard to the former, there still seems to be uncertainty in the psychometric literature about information criteria. For example, many authors continue to refer to “the” DIC (or WAIC or other) value of a model without realizing that multiple values exist for the same model. And the <xref ref-type="bibr" rid="bid__47">Zhang et al. (2019)</xref> simulation results lead them to recommend a “joint” criterion for multilevel IRT models, which differs from the <xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref> marginal recommendation. We consider these issues further below. With regard to complex psychometric models, the nuances of information criterion computation grow with model complexity. We will discuss computational problems that are easy to overlook in psychometric models of ordinal data and of two-level data.</p>
<p>In the pages below, we first define a modeling framework and review the work of <xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref>. We then consider how related problems can manifest themselves in other models, both through imprecision in numerical approximations and through heterogeneity in two-level models. Finally, we conclude with recommendations for use of information criteria in practice and with general remarks about model comparison.</p></sec>
<sec id="s2" sec-type="body"><title>Theoretical Background</title>
<p>We start by introducing a simple structural equation model that can be extended to the models that we consider later. Say that we observe <italic>N</italic> individuals who each report <italic>p</italic> continuous variables. We can model the data from individual <italic>i</italic>, <italic><bold>y</bold><sub>i</sub></italic>, as</p>
<p><disp-formula id="xd1"><label>1</label><mml:math id="xa-1" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mo>+</mml:mo><mml:mo mathvariant="bold">Λ</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where <bold>Λ</bold> is a <italic>p</italic> × <italic>m</italic> (<italic>m</italic> &lt; <italic>p</italic>) matrix of loadings, <bold>η</bold><sub>i</sub> is an <italic>m</italic> × 1 vector of individual <italic>i</italic>’s latent variables, and the remaining terms are <inline-formula id="asda2"><mml:math id="jo-1" display="inline"><mml:mi>p</mml:mi><mml:mo>×</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> vectors. This equation is sometimes called the <italic>measurement submodel</italic>.</p>
<p>A second equation, the <italic>structural submodel</italic>, allows for regression relationships between latent variables. We write it as</p>
<p><disp-formula id="xd2"><label>2</label><mml:math id="xa-2" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">Β</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ζ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where <bold>α</bold> and <bold>ζ</bold><sub>i</sub> are each <italic>m</italic> × 1, and <italic><bold>B</bold></italic> is <italic>m</italic> × <italic>m</italic>. We additionally require that the diagonal of <italic><bold>B</bold></italic> is <bold>0</bold>. This is because it does not make sense for a latent variable to enter into a regression relationship with itself. The residuals <italic><bold>ϵ</bold><sub>i</sub></italic> and <bold>ζ</bold><sub>i</sub> are then typically assumed to be multivariate normal:</p>
<p><disp-formula id="xd3"><label>3</label><mml:math id="xa-3" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ε</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:msub><mml:mi mathvariant="italic">N</mml:mi><mml:mi mathvariant="italic">p</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="bold">Θ</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd4"><label>4</label><mml:math id="xa-4" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ζ</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:msub><mml:mi mathvariant="italic">N</mml:mi><mml:mi mathvariant="italic">m</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="bold">Ψ</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>which leads to a multivariate normal model of the <italic><bold>y</bold><sub>i</sub></italic>.</p>
<p>Of importance for information criterion computations, there are two different likelihoods that are often used for model estimation. The first likelihood conditions on the η<sub>i</sub> and resembles a multivariate regression model:</p>
<p><disp-formula id="xd5"><label>5</label><mml:math id="xa-5" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>Ν</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mo>+</mml:mo><mml:mo mathvariant="bold">Λ</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mo mathvariant="bold">Θ</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This likelihood is commonly used for Bayesian model estimation, where we have an additional posterior distribution of latent variables from which we draw samples. For some complex models with, e.g., ordinal variables or interactions between latent variables, this likelihood greatly simplifies model estimation (e.g., <xref ref-type="bibr" rid="bid__15">Lee et al., 2007</xref>). Use of the conditional likelihood in Bayesian estimation is a type of <italic>data augmentation</italic> (<xref ref-type="bibr" rid="bid__38">Tanner &amp; Wong, 1987</xref>), as it “augments” the observed data with draws of the latent variables.</p>
<p>The second likelihood involves integrating the latent variables out of the model (or we could say <italic>marginalizing over the latent variables</italic>), yielding</p>
<p><disp-formula id="xd6"><label>6</label><mml:math id="xa-6" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>N</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mo>+</mml:mo><mml:mo mathvariant="bold">Λ</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mo>,</mml:mo><mml:mo mathvariant="bold">Λ</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold">Ψ</mml:mi><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mo mathvariant="bold">Λ</mml:mo><mml:mo>′</mml:mo></mml:msup><mml:mo>+</mml:mo><mml:mo mathvariant="bold">Θ</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
    <p>where we require (<italic><bold>I</bold></italic> — <italic><bold>B</bold></italic>) to be invertible. This likelihood is used for frequentist estimation via maximum likelihood and least squares methods.</p>
<p>The conditional likelihood from <xref ref-type="disp-formula" rid="xd5">(5)</xref> is most often used for Bayesian model estimation, especially in Gibbs sampling contexts. But either of the above likelihoods could be used for Bayesian estimation via MCMC, and they will yield the same posterior distributions of model parameters. Importantly, though, the two likelihoods will yield different information criterion values for the same model. We further consider this issue in the next section.</p></sec>
<sec id="s3" sec-type="body"><title>Information Criteria</title>
    <p>Popular information criteria like DIC and WAIC involve evaluations of the model likelihood at posterior parameter values sampled via MCMC. Their general aim is to approximate the model’s ability to predict new (“out of sample”) observations, where models with better out-of-sample predictions are selected over models with worse out-of-sample predictions. These ideas are highly related to the “expected cross-validation index” that was previously developed for SEM (<xref ref-type="bibr" rid="bid__4">Browne &amp; Cudeck, 1989</xref>; <xref ref-type="bibr" rid="bid__7">Cudeck &amp; Browne, 1983</xref>).</p>
<p>All of the criteria that we consider approximate out-of-sample prediction using only in-sample data. This is accomplished by first evaluating the likelihood of the observed data at the posterior samples, which represents the model’s in-sample predictive accuracy. We know that this predictive accuracy will be better than out-of-sample predictive accuracy. Consequently, the information criteria involve an additional penalty term for model complexity. This penalty term is sometimes called the <italic>optimism</italic> or <italic>effective number of parameters</italic>. We consider specifics below.</p>
<sec id="s3.1" sec-type="body"><title>DIC</title>
<p>DIC is the simplest of the criteria considered here. Assuming a generic parameter vector <italic><bold>θ</bold></italic> and <italic>S</italic> posterior samples, it can be written as</p>
<p><disp-formula id="xd7"><label>7</label><mml:math id="xa-7" display="block"><mml:mrow><mml:mtext>DIC</mml:mtext><mml:mo>=</mml:mo><mml:mo>−</mml:mo><mml:mn>2</mml:mn><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>¯</mml:mo></mml:mover><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mi>D</mml:mi></mml:msub></mml:mrow></mml:math></disp-formula></p>
    <p><disp-formula id="xd8"><label>8</label><mml:math id="xa-8" display="block"><mml:mrow><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>S</mml:mi></mml:munderover><mml:mrow><mml:mo>−</mml:mo><mml:mn>2</mml:mn><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mstyle></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo stretchy="false">|</mml:mo><mml:mover accent="true"><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo>¯</mml:mo></mml:mover><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
    <p>where <italic>p<sub>D</sub></italic> is the effective number of parameters, <inline-formula id="mir2.5"><mml:math id="jo-2.5" display="inline"><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>s</mml:mi></mml:msup></mml:math></inline-formula> is the <italic>s</italic>th posterior sample, and <inline-formula id="mir2"><mml:math id="jo-2" display="inline"><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover></mml:math></inline-formula> is the posterior mean (or another posterior measure of central tendency). Notice that the in-sample predictive accuracy log <italic>p</italic>(<italic><bold>y</bold></italic> | <inline-formula id="mir3"><mml:math id="jo-3" display="inline"><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover></mml:math></inline-formula>) is multiplied by −2 and that <italic>p<sub>D</sub></italic> is generally positive. So smaller DIC values are favored over larger values.</p>
<p><xref ref-type="bibr" rid="bid__6">Celeux et al. (2006)</xref> considered many variations of DIC, some of which replace the 2log <italic>p</italic>(<italic><bold>y</bold></italic> | <inline-formula id="mir4"><mml:math id="jo-4" display="inline"><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover></mml:math></inline-formula>) term with other measures of in-sample predictive accuracy. <xref ref-type="bibr" rid="bid__8">Du et al. (2023)</xref> studied some of these variations as applied to multilevel models and found evidence that they can lead to improved model selection decisions over the original DIC. While we do not describe all the variations here, we do note that they can lead to additional confusion about what DIC metric is being reported in a given application.</p></sec>
<sec id="s3.2" sec-type="body"><title>WAIC</title>
<p>WAIC is a “fully Bayesian” information criterion that utilizes pointwise predictive accuracy terms, averaging them over the entire posterior distribution. This is in contrast to DIC, which uses the overall model likelihood to approximate predictive accuracy only at the posterior mean. WAIC assumes that each datapoint is independent of the others, which immediately holds for simple linear models but becomes more complicated for mixed models and for psychometric models. Assuming <italic>N</italic> independent datapoints, we can write WAIC as</p>
<p><disp-formula id="xd9"><label>9</label><mml:math id="xa-9" display="block"><mml:mrow><mml:mtext>WAIC</mml:mtext><mml:mo>=</mml:mo><mml:mo>−</mml:mo><mml:mn>2</mml:mn><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mi>log</mml:mi></mml:mrow></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>S</mml:mi></mml:munderover><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mi>p</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd10"><label>10</label><mml:math id="xa-10" display="block"><mml:mrow><mml:mi>p</mml:mi><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mrow><mml:mtext>Var</mml:mtext></mml:mrow><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:mo stretchy="true">(</mml:mo><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd11"><label>11</label><mml:math id="xa-11" display="block"><mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:mfrac></mml:mrow></mml:mstyle><mml:msup><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>S</mml:mi></mml:munderover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>−</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>S</mml:mi></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover><mml:mo>∑</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>S</mml:mi></mml:munderover><mml:mrow><mml:mi>log</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">θ</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mstyle><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mstyle></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:math></disp-formula></p></sec>
    <sec id="s3.3" sec-type="body"><title>PSIS-LOO</title>
<p><xref ref-type="bibr" rid="bid__45">Watanabe (2010)</xref> showed that WAIC is asymptotically equivalent to leave-one-out cross-validation, whereby we hold out each datapoint, estimate the model, and compute the likelihood of the held-out datapoint given the estimated parameters. <xref ref-type="bibr" rid="bid__41">Vehtari et al. (2017)</xref> propose a Pareto-smoothed importance sampling (PSIS) method for directly approximating leave-one-out cross-validation, as opposed to indirectly approximating leave-one-out cross-validation via WAIC. This method, PSIS-LOO, is implemented in the R package <italic>loo</italic> (<xref ref-type="bibr" rid="bid__40">Vehtari et al., 2024</xref>) and is commonly obtained alongside WAIC. Just like WAIC, the PSIS-LOO metric is based on the pointwise likelihoods evaluated at each posterior sample. The algorithm underlying PSIS-LOO produces diagnostic metrics (“Pareto <italic>k</italic>” metrics) that can signal problems with the approximation as well as influential observations in the dataset. The availability of diagnostics leads us to prefer PSIS-LOO over WAIC, though the two metrics are usually very similar to one another for the psychometric models considered in this paper.</p></sec>
<sec id="s3.4"><title>Choice of Likelihood</title>
<p>The above criteria involve either the overall model likelihood <italic>p</italic>(<italic><bold>y</bold></italic> | <bold>θ</bold>) or the pointwise likelihoods <italic>p</italic>(<italic><bold>y</bold><sub>i</sub></italic> | <bold>θ</bold>). As seen by Equations <xref ref-type="disp-formula" rid="xd5">(5)</xref> and <xref ref-type="disp-formula" rid="xd6">(6)</xref>, a psychometric model with latent variables has multiple likelihoods. The specific choice of likelihood influences the values of the information criteria and the type of predictive generalization that the information criteria approximate (see <xref ref-type="bibr" rid="bid__23">Merkle et al., 2019</xref>).</p>
<p>MCMC estimation of Bayesian psychometric models typically involves sampling latent variables for each person. By conditioning on these latent variables, we usually attain independence between individual observations within a person. The likelihood <italic>p</italic>(<italic><bold>y</bold><sub>i</sub></italic> | <bold>θ</bold>) is then based on Equation <xref ref-type="disp-formula" rid="xd5">(5)</xref>, where we use conditional independence to multiply across the univariate likelihoods <italic>p</italic>(<italic>y<sub>ij</sub></italic> | <bold>θ</bold>). When we compute information criteria using this conditional likelihood, we approximate the model’s ability to predict new data from the same people in the original dataset because we have conditioned on the latent variables η<sub>i</sub>. <xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref> referred to this as “leave one unit out” cross-validation.</p>
<p>Frequentists typically use a marginal likelihood for model estimation, where the latent variables have been integrated out of the model. In this case, the likelihood <italic>p</italic>(<italic><bold>y</bold><sub>i</sub></italic> | <bold>θ</bold>) is based on Equation <xref ref-type="disp-formula" rid="xd6">(6)</xref>. For IRT models and some others, the marginal criteria require quadrature or related numerical methods because an analytic expression like Equation <xref ref-type="disp-formula" rid="xd6">(6)</xref> cannot be obtained. When we compute information criteria using the marginal likelihood, we approximate the model’s ability to predict new data from new people who were not in the original dataset. <xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref> referred to this as “leave one cluster out” cross-validation.</p>
    <p>In typical psychometric applications, there are many reasons to prefer marginal information criteria. First, we typically view individuals as exchangeable, and we wish to generalize beyond the people that were observed in our data (see <xref ref-type="bibr" rid="bid__21.75">Merkle, 2026b</xref>). Marginal information criteria approximate such generalization. Second, <xref ref-type="bibr" rid="bid__26">Millar (2018)</xref> discusses the fact that conditional forms of WAIC may not be asymptotically equivalent to leave-one-out cross-validation because the number of model parameters grows with the sample size. So there is less theoretical justification for conditional information criteria, as compared to marginal information criteria. Finally, <xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref> showed that marginal information criteria have less Monte Carlo error than conditional information criteria because the marginal likelihoods involve many fewer parameters. This means that we can obtain precise values of marginal information criteria using fewer posterior samples, as compared to conditional information criteria.</p>
<p>Despite these advantages, marginal information criteria are not always used in practice. This is especially true because traditional MCMC estimation methods sample the latent variables and rely on a conditional likelihood. So researchers will immediately think of the conditional likelihood as the one to use for information criteria, and some software like JAGS will automatically compute DIC for a specified model (which usually relies on the conditional likelihood). Additionally, the marginal likelihood is often more difficult to compute than the conditional likelihood. This can be frustrating to researchers who estimate a Bayesian model with a conditional likelihood, only to find out that they need further machinery to obtain marginal information criteria.</p></sec>
 <sec id="s3.5" sec-type="body"><title>Comparison of Marginal and Conditional Criteria</title>
<p>If the conditional and marginal information criteria always led to the same conclusions, then their distinction could be ignored. On this point, <xref ref-type="bibr" rid="bid__23">Merkle et al. (2019)</xref> presented theoretical results that the conditional and marginal criteria differ, along with two examples where the criteria selected different psychometric models. The conditional criteria generally favored more complex models than marginal criteria, which makes sense when we consider the types of cross-validation that they represent. For example, we already stated that conditional criteria are related to predicting new data from people that we already observed. From this standpoint, we can afford additional model complexity that tunes parameters to the people in the observed data. In contrast, a marginal criterion needs to guard against overfitting to the people in the observed data because it is related to predicting unseen people.</p></sec>
<sec id="s3.6" sec-type="body"><title>On Use of Joint Likelihoods</title>
<p>Other researchers sometimes use information criteria that are not fully marginal or conditional. Most related to the current paper, <xref ref-type="bibr" rid="bid__47">Zhang et al. (2019)</xref> recommend a “joint” form of DIC in a two-level IRT model of, e.g., students nested in schools. They do not consider WAIC or PSIS-LOO, and we have not seen joint versions of WAIC or PSIS-LOO used in practice.</p>
<p>Using our notation, the joint DIC involves the likelihood <italic>p</italic>(<bold>y</bold><sub>i</sub>, <bold>η</bold><sub>i</sub> | <bold>θ</bold>) = <italic>p</italic>(<italic><bold>y</bold><sub>i</sub></italic> | <bold>η</bold><sub>i</sub>, <bold>θ</bold>)<italic>p</italic>(<bold>η</bold><sub>i</sub> | <bold>θ</bold>), where <bold>θ</bold> would contain all model parameters except the latent variables η<sub>i</sub>. This joint likelihood is typically easier to compute than the marginal likelihood, and it appears to have been proposed because it is convenient (see <xref ref-type="bibr" rid="bid__6">Celeux et al., 2006</xref>). It resembles a one-node quadrature approximation of the marginal likelihood (e.g., <xref ref-type="bibr" rid="bid__32">Rabe-Hesketh et al., 2005</xref>), where the specific node varies for each posterior sample and where the quadrature weights do not sum to 1. <xref ref-type="bibr" rid="bid__47">Zhang et al. (2019)</xref> describe some simulations where a joint DIC selects the true model more often than a marginal DIC.</p>
<p>From our point of view, use of the joint likelihood for information criteria is theoretically problematic because it does not have a clear relationship to cross-validation. That is, it is unclear what it means to hold out a model parameter (the <bold>η</bold><sub>i</sub>) in addition to <italic><bold>y</bold><sub>i</sub></italic> when considering a model’s predictive accuracy. Inclusion of the <bold>η</bold><sub>i</sub> in the likelihood will also lead to a large amount of noise in the resulting information criteria, a point that <xref ref-type="bibr" rid="bid__47">Zhang et al. (2019)</xref> mention in their paper. These considerations lead us to maintain the default recommendation of fully marginal information criteria, because these criteria match the type of generalization in which researchers are usually interested. We further consider these recommendations in the <xref ref-type="sec" rid="s6">General Discussion</xref>.</p>
<p>While our information criterion recommendations differ from those of <xref ref-type="bibr" rid="bid__47">Zhang et al. (2019)</xref>, we definitely agree with them that marginal forms of information criteria can be computationally problematic. In addition to long computation times, we describe here two examples where likelihood approximations and model misspecifications can lead to unexpected behavior of the marginal criteria. In the first example, we consider a model of ordinal variables that requires numerical approximation of the marginal likelihood. We show how crude numerical approximations lead to imprecise information criteria. In the second example, we consider two-level SEM and show how information criteria are impacted by heterogeneity. Beyond raising awareness for these specific problems, we hope that these examples make it clear that use of information criteria requires some care.</p></sec></sec>
<sec id="s4" sec-type="body"><title>Example 1: Impact of Numerical Approximation</title>
<p>Numerical approximations of marginal likelihoods are often needed in psychometric models of ordinal data. Structural equation models of ordinal data can be motivated by continuous, latent data vectors <inline-formula id="mir5"><mml:math id="jo-5" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> underlying the ordinal data <italic><bold>y</bold><sub>i</sub></italic>. We take the structural equation model from Equations <xref ref-type="disp-formula" rid="xd1">(1)</xref> and <xref ref-type="disp-formula" rid="xd2">(2)</xref> and place it on the <inline-formula id="mir6"><mml:math id="jo-6" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula>, with threshold parameters <bold>τ</bold> that chop entries of <inline-formula id="mir7"><mml:math id="jo-7" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> into ordered categories. For example, in the case of an ordinal variable <italic>y<sub>ij</sub></italic> with four categories, we would have:</p>
<p><disp-formula><mml:math id="as1" display="block"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mspace width="0.5em"/><mml:mtext>=</mml:mtext><mml:mspace width="0.5em"/><mml:mtext>1</mml:mtext><mml:mspace width="0.5em"/><mml:mtext>if</mml:mtext><mml:mspace width="0.5em"/><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mtext>&lt;</mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mspace width="0.5em"/></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>2 if </mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&lt;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>3 if </mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&lt;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>4 if </mml:mtext><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mtext>&gt;</mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math>
</disp-formula></p>
<p>where τ<sub><italic>j</italic>1</sub> &lt; τ<sub><italic>j</italic>2</sub> &lt; τ<sub><italic>j</italic>3</sub>.</p>
    <sec id="s4.1" sec-type="body"><title>Model Likelihoods</title>
<p>Traditional MCMC methods for estimating this model involve sampling either the <inline-formula id="mir8"><mml:math id="jo-8" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> or the <bold>η</bold><sub>i</sub>, both of which are a form of data augmentation. In the former case, we can condition on the sampled <inline-formula id="mir9"><mml:math id="jo-9" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> and pretend like we have continuous data when sampling the other model parameters. In the latter case, we achieve independence of the entries of <italic><bold>y</bold><sub>i</sub></italic> conditioned on <bold>η</bold><sub>i</sub>. Then we can evaluate the (conditional) likelihood of the <italic><bold>y</bold><sub>i</sub></italic> via univariate normal CDFs, which is computationally easy.</p>
<p>Neither of these estimation methods is optimal from an information criterion point of view. This is because the marginal likelihood of the <italic><bold>y</bold><sub>i</sub></italic> is marginal over both the <inline-formula id="mir10"><mml:math id="jo-10" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> and the <bold>η</bold><sub>i</sub>. This marginal likelihood can be computed with the CDF of the multivariate normal distribution, but its evaluation is generally too slow to be useful for MCMC estimation. We can instead estimate the model (draw posterior samples) by sampling the <inline-formula id="mir11"><mml:math id="jo-11" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> or the <bold>η</bold><sub>i</sub>, then compute the fully marginal likelihood of the <italic><bold>y</bold><sub>i</sub></italic> for each posterior sample after estimation.</p>
<p>Even if the fully marginal likelihood is computed after estimation, it is still computationally expensive. This is especially the case for metrics like WAIC and PSIS-LOO that require pointwise likelihoods. For example, say that we have a dataset of 300 individuals, and we save 1,000 posterior samples for each of three chains. If it takes one second to evaluate the pointwise log-likelihoods for one posterior sample, then it will take 50 minutes to evaluate the log-likelihoods of all 3,000 posterior samples. It is clear that we need to be as economical as possible, while still doing accurate computations. There exist multiple options for approximating the marginal likelihood, and we consider two popular options below.</p>
<sec id="s4.1.1" sec-type="body"><title>Adapted Gauss-Hermite Quadrature</title>
    <p>Because we are approximating the model likelihood after model estimation, we have access to samples of the latent variables <bold>η</bold><sub>i</sub>. This allows us to make informed choices about quadrature nodes for each case <italic>i</italic> (see <xref ref-type="bibr" rid="bid__23">Merkle et al., 2019</xref>; <xref ref-type="bibr" rid="bid__32">Rabe-Hesketh et al., 2005</xref>) so that the procedure is “already adapted” instead of “adaptive”. As the number of quadrature nodes per dimension increases, Gauss-Hermite quadrature is often treated as the gold standard approximation. But large numbers of quadrature nodes are computationally expensive, even when the nodes are already adapted.</p></sec>
<sec id="s4.1.2" sec-type="body"><title>Importance Sampling of <inline-formula id="mir12"><mml:math id="jo-12" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula></title>
<p>There exist sampling-based methods for approximating the truncated multivariate normal likelihood, the most popular of which is the GHK algorithm (e.g., <xref ref-type="bibr" rid="bid__12">Hajivassiliou &amp; McFadden, 1998</xref>). This involves transforming the multivariate normal to have mean zero and identity covariance matrix, then generating samples of <inline-formula id="mir13"><mml:math id="jo-13" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula> via a series of truncated univariate distributions. For the purposes of information criteria, we can additionally obtain an estimate of the marginal likelihood of <italic><bold>y</bold><sub>i</sub></italic> by averaging across importance sampling weights obtained during sampling. Larger numbers of importance samples lead to more accurate approximations, similarly to larger numbers of quadrature nodes. The <italic>tmvnsim</italic> package (<xref ref-type="bibr" rid="bid__2">Bhattacjarjee, 2016</xref>) provides an implementation of the GHK method in R.</p>
<p>For most traditional SEMs, the dimension of <bold>η</bold><sub>i</sub> is less than the dimension of the <inline-formula id="mir14"><mml:math id="jo-14" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula>, suggesting that quadrature is advantageous over importance sampling and other methods involving truncated multivariate normal distributions (because quadrature integrates over η, and others integrate over <italic>y</italic><sup>∗</sup>). But all the methods involve a tradeoff between computational accuracy and speed, with this tradeoff potentially having implications for information criteria. This is examined in the next section.</p></sec></sec>
<sec id="s4.2" sec-type="body"><title>Illustration</title>
<p>Here, we study how crude approximations of the marginal log-likelihood influence the resulting information criterion values. We fit a model to an ordinal dataset, approximate the pointwise contributions to the marginal log-likelihood in various ways, and compare the resulting information criteria to a gold standard.</p>
<sec id="s4.2.1" sec-type="body"><title>Method</title>
<p>We fit an ordinal factor analysis model to items from the Nerdy Personality Attributes Scale (<xref ref-type="bibr" rid="bid__28">Open Source Psychometrics Project, 2016</xref>), which was previously used by <xref ref-type="bibr" rid="bid__34">Schneider et al. (2020)</xref> to study frequentist model comparison statistics. The scale consists of 26 five-point Likert items that are designed to measure a person’s “nerdiness.” We focus here on 1018 responses to the 6-item science subscale. An example item is “I prefer academic success to social success,” with the lowest option being “disagree” and the highest option being “agree.”</p>
<p>We fit a 1-factor model to the data, which, from an IRT point of view, can be called a graded response model with probit link function. We place the model from <xref ref-type="disp-formula" rid="xd6">Equation (6)</xref> on the latent responses <inline-formula id="mir15"><mml:math id="jo-15" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mo>∗</mml:mo></mml:msubsup></mml:math></inline-formula>, with the individual <italic>y<sub>ij</sub></italic> being obtained via:</p>
    <p><disp-formula><mml:math id="jo-16" display="block"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>1 if </mml:mtext><mml:mspace width="0.25em"/><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>2 if </mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&lt;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>3 if </mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&lt;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>4 if </mml:mtext><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>&lt;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mo>&lt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>5 if </mml:mtext><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mo>*</mml:mo></mml:msubsup><mml:mspace width="0.25em"/><mml:mo>&gt;</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where τ<sub><italic>j</italic>1</sub> &lt; τ<sub><italic>j</italic>2</sub> &lt; τ<sub><italic>j</italic>3</sub> &lt; τ<sub><italic>j</italic>4</sub>. Prior distributions were intended to be mildly informative and were set as:</p>
    <p><disp-formula id="assd1"><mml:math id="jo-17" display="block"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mo>λ</mml:mo><mml:mi>j</mml:mi></mml:msub><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mspace width="0.25em"/><mml:mtext>Normal</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.4</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.05em"/><mml:mspace width="0.05em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>6</mml:mn></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>ψ</mml:mo><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow/><mml:mspace width="0.5em"/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Gamma</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mo>τ</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>⁢</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mspace width="0.25em"/><mml:mtext>Normal</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.05em"/><mml:mspace width="0.05em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>6</mml:mn></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula><disp-formula id="assd1.5"><mml:math id="jo-17.5" display="block"><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mi>log</mml:mi><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>⁢</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mi>τ</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Normal</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.05em"/><mml:mspace width="0.05em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>6</mml:mn><mml:mo>;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>4.</mml:mn></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>To identify the model, we fixed λ<sub>1</sub> = 1 and θ<sub>j</sub> = 1 for <italic>j</italic> = 1<italic>,...,</italic> 6. For each of three chains, we used <italic>blavaan</italic> to obtain 2,000 samples following 500 burnin samples. This is more samples than we need for precise estimates of posterior means, but the information criteria generally have more Monte Carlo error than posterior means. In past work, if it takes <italic>S</italic> posterior samples for stable posterior means, then we take 2<italic>S</italic> posterior samples for the information criteria. A more formal examination can also be useful, say by monitoring the model log-likelihood during MCMC and examining its effective sample size.</p>
<p>Following model estimation, we computed DIC and PSIS-LOO using a variety of marginal likelihood approximations. Our gold standard marginal log-likelihood approximation is an adapted quadrature method with a large number of nodes. For the particular model that we consider here, 8 quadrature nodes per dimension was sufficient (further numbers of nodes left the log-likelihoods virtually unchanged). We compared the gold standard to GHK sampling with 5, 10, 20, and 50 samples per case, and to adapted quadrature with 1, 2, and 4 nodes.</p></sec>
<sec id="s4.2.2" sec-type="body"><title>Results</title>
    <p>It is difficult to make exact statements about the computation times of the marginal likelihood approximations because the code that we used (see <xref ref-type="bibr" rid="bid__21.25">Merkle, 2026a</xref>) was not optimized for parallelization or for all available computational shortcuts. But to provide a rough idea of timings, each set of marginal likelihood computations took at least 5 minutes per chain on a mid-range laptop, with some taking closer to an hour per chain. GHK sampling was faster than quadrature, likely because the GHK code had better optimizations. Cruder approximations (fewer quadrature points or importance samples) provided speedups of up to 30%.</p>
<p><xref ref-type="fig" rid="f1">Figure 1</xref> shows our primary results, with information criteria on the y-axes and number of quadrature points or importance samples on the x-axes. The columns show the adapted quadrature and GHK methods, and the rows show DIC and PSIS-LOO. The dashed line is the gold standard value with 8 quadrature points, and shaded areas represent 1 standard error around the criterion value (note that standard errors are not available for DIC). We see that the quadrature method generally yields information criterion values that are close to the gold standard. For GHK sampling, the criterion values are slightly larger when the approximation is crude. These differences look small, with the dashed line being well within one standard error of the PSIS-LOO values. But information criteria are generally correlated across models, so that differences between model criteria will have much smaller standard errors. For example, it is possible to compare the gold standard PSIS-LOO to other PSIS-LOO metrics, pretending like the metrics were coming from different models. In doing so, we find PSIS-LOO differences of 10 standard errors or more.</p>
    <fig id="f1" position="anchor" fig-type="figure" orientation="portrait"><label>Figure 1</label><caption><title>Information Criteria of Ordinal Factor Analysis Model</title><p><italic>Note</italic>. Criteria are computed using marginal likelihoods approximated by adapted quadrature and by GHK sampling with differing numbers of nodes/samples. Shaded regions represent ±1 standard error.</p></caption><graphic xlink:href="meth.20361-f1.png" position="anchor" orientation="portrait"/></fig>
    <p><xref ref-type="fig" rid="f2">Figure 2</xref> is arranged similarly to <xref ref-type="fig" rid="f1">Figure 1</xref>, except it shows the estimated effective number of parameters instead of the information criteria. We see here that, under crude approximations of the marginal likelihood, the effective number of parameters is usually estimated to be too large. The exception is the upper right panel, which shows DIC with GHK sampling. There, effective number of parameters is too low until we obtain 40 samples or more for the approximation.</p><fig id="f2" position="anchor" fig-type="figure" orientation="portrait"><label>Figure 2</label><caption><title>Estimated Effective Numbers of Parameters Associated with DIC and PSIS-LOO</title><p><italic>Note</italic>. Effective numbers of parameters are computed using marginal likelihoods approximated by adapted quadrature and by GHK sampling with differing numbers of nodes/samples. Shaded regions represent ±1 standard error.</p></caption><graphic xlink:href="meth.20361-f2.png" position="anchor" orientation="portrait"/></fig></sec>
<sec id="s4.2.3" sec-type="body"><title>Discussion</title>
<p>For the model considered here, crude approximations of the marginal log-likelihood led to information criteria that were sometimes too large. The GHK method with 5 samples is the crudest approximation here, and we found that its influence on effective number of parameter computations differed across DIC and PSIS-LOO (leading to a low value for DIC and a high value for PSIS-LOO). We believe that this is because, for PSIS-LOO, the effective number of parameter computations involve variances of pointwise log-likelihoods, so that imprecision from the crude integral approximations leads to upward bias. This result is related to the work of <xref ref-type="bibr" rid="bid__39">Timonen et al. (2023)</xref>, who proposed importance weights that correct for crude numerical approximations in Ordinary Differential Equation models. DIC, on the other hand, involves averages of the full model log-likelihood that are differently affected by crude approximations.</p>
    <p>In some cases, the marginal likelihood approximations will not influence model selection decisions. This is because the extra approximation imprecision will be small compared to the differences in the models’ out-of-sample predictive accuracy. More caution may be warranted in situations where we compare models with different numbers of latent variables. For example, if one model involves a 1-dimensional latent variable and another model involves a three-dimensional latent variable, we may expect the second model to suffer more from the integral approximation. And in situations where models have the same number of latent variables, we may still be cautious about selecting one model when the information criterion values are close (also see <xref ref-type="bibr" rid="bid__25">Merkle et al., 2016</xref>). In the next section, we consider how model misspecifications can lead to unexpected behavior of information criteria.</p><?figure f1?><?figure f2?></sec></sec></sec>
<sec id="s5" sec-type="body"><title>Example 2: Heterogeneity in Two-Level Models</title>
<p>Two-level structural equation models lead to additional conditional/marginal distinctions in information criteria because they involve latent variables at multiple levels. These models are traditionally applied to datasets of students nested within schools or within countries. We observe multiple test scores and/or other variables per student, which leads us to specify student-level latent variables as in traditional SEM. But students within the same school are correlated with one another, and we want our model to additionally account for these correlations. We review this modeling framework first before considering problems with information criteria.</p>
<sec id="s5.1"><title>Model</title>
<p>We now define the vector <italic><bold>y</bold><sub>ik</sub></italic> to be the data of Level 1 unit <italic>i</italic> within Level 2 unit <italic>k</italic>. In traditional applications, <italic>i</italic> would be student and <italic>k</italic> would be school, and we generally refer to students and schools below because it helps with intuition. Mixed modeling researchers may view this as a three-level dataset where individual observations are nested within students, and students are nested within schools. See <xref ref-type="bibr" rid="bid__35">Skrondal and Rabe-Hesketh (2004)</xref> for further discussion of these different viewpoints.</p>
<p>Specification of a two-level SEM with random intercepts can look very similar to a one-level SEM. We start out writing</p>
<p><disp-formula id="xd12"><label>12</label><mml:math id="xa-12" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mo mathvariant="bold">Λ</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold">ε</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd13"><label>13</label><mml:math id="xa-13" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">B</mml:mi><mml:msub><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ζ</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>which is nearly the same as the traditional model from Equations <xref ref-type="disp-formula" rid="xd1">(1)</xref> and <xref ref-type="disp-formula" rid="xd2">(2)</xref>. The only differences are that, (i) our intercept <bold>ν</bold> has a <italic>k</italic> subscript, and (ii) some other vectors have <italic>ik</italic> subscripts because we are designating the model for student <italic>i</italic> in school <italic>k</italic>.</p>
<p>The <bold>ν</bold><sub>k</sub> term implies that the intercept is unique to each school <italic>k</italic>, which is what leads us to call this a “two-level” SEM. We complete the model by considering how to handle the <bold>ν</bold><sub>k</sub>. If we come from a mixed modeling background, the natural thing to do is place a multivariate normal hyperdistribution on the <bold>ν</bold><sub>k</sub>, i.e.,</p>
<p><disp-formula id="xd14"><label>14</label><mml:math id="xa-14" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>Ν</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This can add many parameters to the model, especially when each individual has many observed variables. For example, if each individual has <italic>p</italic> = 9 observed variables, then <bold>Σ</bold><sub>ν</sub> contains 45 free parameters.</p>
<p>Because an unrestricted multivariate normal can lead to many extra parameters, and also because we are already using latent variables, we could elect to place a second SEM on the <bold>ν</bold><sub>k</sub>. This typically involves specifying school latent variables that are predictive of individual entries of <bold>ν</bold><sub>k</sub>. We basically repeat our SEM, modeling <bold>ν</bold><sub>k</sub> instead of <italic><bold>y</bold><sub>ik</sub></italic>:</p>
<p><disp-formula id="xd15"><label>15</label><mml:math id="xa-15" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msup><mml:mo mathvariant="bold">Λ</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msubsup><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="bold">ε</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd16"><label>16</label><mml:math id="xa-16" display="block"><mml:mrow><mml:msubsup><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msubsup><mml:mi mathvariant="bold-italic">η</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi mathvariant="bold-italic">ζ</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where the (<italic>c</italic>) superscripts distinguish the school-level matrices from the student-level matrices. The mean intercept across schools, <bold>ν</bold><sub>0</sub>, does not have a superscript because it is similar to the intercept in a traditional, one-level SEM. And while we do not consider it here, the model may include school-level observed variables that do not exist for individual students. In that case, the <bold>ν</bold><sub>k</sub> vector is augmented with the observed variables from school <italic>k</italic> (see <xref ref-type="bibr" rid="bid__33">Rosseel, 2021</xref>).</p></sec>
<sec id="s5.2"><title>Conditional and Marginal Distributions</title>
<p>The model has student-level latent variables as well as school-level latent variables. Traditional model estimation via MCMC would proceed by sampling all the latent variables, which typically provides conditional independence between observed variables. This leads to a univariate conditional model likelihood that is relatively simple to evaluate, but it also has implications for information criteria. Depending on what latent variables are (and are not) marginalized out of the model, our information criteria approximate out-of-sample prediction at different levels (also see <xref ref-type="bibr" rid="bid__47">Zhang et al., 2019</xref>):</p>
<list list-type="bullet">
<list-item>
<p>No marginalization (fully conditional likelihood): Predict new observations from the same students that were in the original data (who are also in the same schools).</p></list-item>
<list-item>
<p>Marginalize over student latent variables (partially marginal likelihood): Predict new observations from new students in the same schools that were in the original data.</p></list-item>
<list-item>
<p>Marginalize over both student and school latent variables (fully marginal likelihood): Predict new observations from new students in new schools.</p></list-item>
</list>
<p>Note that our terminology refers only to marginalization of latent variables and should not be confused with the marginal likelihood used for Bayes factor computations. The latter marginalizes over all model parameters, including the parameters that would be viewed as fixed in a frequentist context.</p>
<p>In most psychometric modeling situations, researchers do not care about the specific students and/or schools in the data. So, for computing information criteria, the fully marginal likelihood would be the recommended default. Exceptions may be studies where researchers are assessing the specific students in the dataset, or school accountability studies where researchers care about the specific schools in the dataset. The former may warrant a fully conditional likelihood, and the latter may warrant a partially marginal likelihood.</p>
<p>Unfortunately, the two-level, fully marginal likelihoods involved in the default recommendation can be computationally difficult. The fully marginal likelihood accounts for the fact that students within a school are correlated, which typically leads to a high-dimensional multivariate normal distribution. As a specific example, imagine that we observe 30 students in each school, and each student provides 5 observed variables. Then our fully marginal likelihood above will involve a 150-dimensional normal distribution. It is difficult to evaluate this high-dimensional likelihood, which leads to excessively slow computations.</p>
<p>Fortunately, a good deal of attention has been devoted to this problem, and a large amount of progress has been made (e.g., <xref ref-type="bibr" rid="bid__9">du Toit &amp; du Toit, 2008</xref>; <xref ref-type="bibr" rid="bid__21">McDonald, 1993</xref>; <xref ref-type="bibr" rid="bid__33">Rosseel, 2021</xref>). This is because the fully marginal likelihood is used for frequentist model estimation, and improved computations can lead to faster frequentist estimation. In the previous work, researchers took advantage of the structure of the model covariance matrix to evaluate the likelihood using matrices of low dimension. Let <italic>n<sub>k</sub></italic> be the number of students in school <italic>k</italic>, each of whom provides <italic>p</italic> observed variables. The previous authors have derived expressions that allow us to evaluate the likelihood using matrices of dimension <italic>p</italic>, instead of using matrices of dimension <italic>n<sub>k</sub>p</italic>. The appendix includes technical detail on computing the fully marginal, pointwise log-likelihood of each Level-2 unit in a two-level SEM with random intercepts.</p></sec>
<sec id="s5.3"><title>Illustration</title>
<p>In the two-level SEM framework, it is possible to observe multiple types of heterogeneity. Here, we consider a situation where the within-cluster covariance matrix differs across clusters. We generate data from a two-level model while manipulating the magnitude of heterogeneity. We then fit a model that ignores this heterogeneity and examine the resulting information criteria.</p>
<sec><title>Method</title> 
   <p>We generated data from a model with 4 observed variables and 200 clusters of size 50. The data-generating model can be written as:</p>
    <p><disp-formula id="xd17"><label>17</label><mml:math id="jo-18" display="block"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>⁢</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">ν</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mo>∼</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnspacing="1em" rowspacing=".2em" columnalign="left left"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>N</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">ν</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>50</mml:mn><mml:mo>;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>100</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mtext>N</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>ν</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>×</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>50</mml:mn><mml:mo>;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>101</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>200</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"/></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math>
</disp-formula></p>
    <p><disp-formula id="xd18"><label>18</label><mml:math id="xa-18" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>∼</mml:mo><mml:mtext>Ν</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mn mathvariant="bold">0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mspace width="0.25em"/><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>...</mml:mn><mml:mo>,</mml:mo><mml:mn>200</mml:mn></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd19"><label>19</label><mml:math id="xa-19" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd20"><label>20</label><mml:math id="xa-20" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.6</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.6</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.6</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.2</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.6</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
<p>where we manipulated the scalar <italic>d</italic> to introduce heterogeneity in the within-cluster covariance matrix. The <italic>d</italic> parameter assumed values from 1 to 4 in steps of 0.5, where a new dataset was generated for each value of <italic>d</italic>. We then fit a traditional two-level model that assumes a single within-cluster covariance matrix for all clusters. In <italic>(b)lavaan</italic>, this model could be specified as</p><preformat position="anchor" orientation="portrait" xml:space="preserve">
model &lt;-'
    level: within
        fw =~ y1 + y2 + y3 + y4
    level: between
        fb =~ y1 + y2 + y3 + y4
'
</preformat><fig id="f3" position="anchor" fig-type="figure" orientation="portrait"><label>Figure 3</label><caption><title>Magnitude of Heterogeneity in Within-Cluster Covariance Matrix (X-Axis) vs. Estimated Effective Number of Parameters (Y-Axis) of DIC and PSIS-LOO.</title></caption><graphic xlink:href="meth.20361-f3.png" position="anchor" orientation="portrait"/></fig>
<p>The prior distributions that we used for model estimation were</p>
    <p><disp-formula id="xd21a"><mml:math id="jo-19" display="block"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">ν</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Normal</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>λ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Normal</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.4</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.5em"/><mml:mspace width="0.5em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>4</mml:mn></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>λ</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Normal</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.4</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.5em"/><mml:mspace width="0.5em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>4</mml:mn></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>ψ</mml:mi></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Gamma</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>ψ</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Gamma</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>θ</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Gamma</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.5em"/><mml:mspace width="0.5em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>4</mml:mn></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>θ</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>Gamma</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>.5</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.5em"/><mml:mspace width="0.5em"/><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>4.</mml:mn></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
    <p>where λ<sub>1</sub> = <inline-formula><mml:math id="jo-20" display="inline"><mml:msubsup><mml:mi>λ</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> = 1 and the ψ and θ parameters are standard deviations, as opposed to variances. The priors on loadings are intended to be mildly informative about the direction of correlation between observed variables (i.e., we expect all observed variables to be positively related). The remaining priors are relatively uninformative, considering the data generating model. For each of three chains, we used <italic>blavaan</italic> to obtain 2,000 samples following 500 burnin samples.</p></sec>
    <sec><title>Results</title><p>Results are shown in <xref ref-type="fig" rid="f3">Figure 3</xref>. The value of <italic>d</italic> is on the x-axis, where 1 represents no heterogeneity and values above 1 represent increasing magnitudes of heterogeneity. The estimated effective numbers of parameters are on the y-axis, panels represent different information criteria, and shaded areas represent ±1 standard error (which is unavailable for DIC). The dashed, horizontal line in each panel is the frequentist parameter count, which is 20 for this model. We see that the DIC effective number of parameters is nearly flat and stays near the frequentist parameter count of 20. This is common behavior under mildly informative priors. But under PSIS-LOO, the model appears to gain parameters as the heterogeneity increases. While it is difficult to see in the figure, the standard error for the PSIS-LOO effective number of parameters nearly doubles from <italic>d</italic> = 1 (no heterogeneity) to <italic>d</italic> = 4. The standard error of the full PSIS-LOO criterion (as opposed to the standard error for effective number of parameters) increases more dramatically, from a value of 282 at <italic>d</italic> = 1 to a value of 1,717 at <italic>d</italic> = 4. And even at <italic>d</italic> = 4, the Pareto <italic>k </italic>diagnostics are all in the range that the <italic>loo</italic> package designates as “good.”</p>
<p>When using the marginal likelihood, PSIS-LOO approximates predictive accuracy of new students in new schools. But when there is large heterogeneity from school to school, it becomes difficult to use the existing data to make predictions for new schools. The inflated effective number of parameters reflects the idea that, under the heterogeneity considered here, we cannot effectively pool information across schools. Our result is similar to the “scaled 8 schools” example of <xref ref-type="bibr" rid="bid__41">Vehtari et al. (2017)</xref>, who considered a related problem in a simple two-level model. They showed that WAIC and PSIS-LOO break down when there is no pooling across schools, that is, when data from one school are not informative about data from other schools. We have essentially illustrated a similar issue in a larger model. A related issue can be observed when estimating two-level models with “within only” variables that appear only at Level 1. See the example in Appendix B.</p>
        <p>It may be tempting to view DIC as advantageous in this example because its effective number of parameters remains similar, regardless of the amount of heterogeneity in the data. But we find the behavior of PSIS-LOO worthwhile because it can be used as a diagnostic that signals model misspecification. The unexpected increase in effective number of parameters can potentially lead researchers to do further posterior checking and to improve the model.</p></sec></sec> </sec>
<sec id="s6" sec-type="body"><title>General Discussion</title>
<p>In this paper, we first reviewed the idea that multiple versions of Bayesian information criteria are available for the same model, depending on how one handles the latent variables and random effects. We recommend use of marginal likelihoods (marginal over latent variables and random effects) as the default because the resulting information criteria approximate predictive generalization to new people and to new Level-2 units, in the two-level case. This type of generalization is usually of primary interest, as compared to generalization to new data from the same people and/or from the same Level-2 units that were in the original dataset. The marginal information criteria also have better theoretical justification and less Monte Carlo error, as compared to conditional information criteria.</p>
<p>We then highlighted the idea that computation of Bayesian information criteria can be more intricate than one would expect. Multiple criteria exist for the same model, numerical approximations that are not involved in model estimation may still influence the information criteria, and heterogeneity in two-level models may lead to unexpected behavior in the information criteria. These results are especially problematic for researchers who view information criteria as simple methods for model comparison that do not require much extra thought. In the sections below, we further consider applied use of information criteria and related methods in psychometrics.</p>
<sec id="s6.1" sec-type="body"><title>Computational Issues</title>
    <p>Marginal information criteria are generally more difficult to compute, as compared to conditional information criteria, because they often require extra computations after model estimation that may not be needed otherwise. Software tools can facilitate some of these extra computations. These include the <italic>blavaan</italic> package, which automatically computes marginal information criteria for the models that it is able to estimate, and the <italic>bleval</italic> package (<xref ref-type="bibr" rid="bid__17">Luo &amp; Dong, 2025</xref>; <xref ref-type="bibr" rid="bid__18">Luo et al., 2026</xref>), which contains a quadrature method that can be flexibly applied to the posterior samples of many Bayesian models. The latter package requires users to specify a function defining the model’s conditional likelihood.</p>
<p>While the above packages make marginal likelihood evaluations easier, the computations will still be intensive and/or infeasible for some models. While this may lead researchers to use conditional or joint likelihoods, we caution that those criteria do not necessarily approximate the generalization that would often be expected of psychometric models. Researchers may instead consider alternative model evaluation methods and computations. For example, the <italic>loo</italic> package authors and their colleagues recently developed a subsampling approach to computing PSIS-LOO for large datasets (<xref ref-type="bibr" rid="bid__19">Magnusson et al., 2020</xref>; <xref ref-type="bibr" rid="bid__20">Magnusson et al., 2025</xref>). This approach does not require us to compute the likelihood for every observation in the dataset so can lead to faster computations, while also providing estimates of additional uncertainty due to subsampling. Alternatively, it is possible to directly carry out a cross-validation analysis by sequentially holding out data, fitting models, and examining predictions of the held-out data via metrics other than log-likelihoods. This may be advantageous when model estimation is fast and marginal likelihood computation is slow. As a third alternative, there might be an encompassing model that contains all candidate models as special cases. We can use the encompassing model to account for our uncertainty in candidate models, and the encompassing model also facilitates other model comparison metrics such as the Bayes factor (e.g., <xref ref-type="bibr" rid="bid__44">Wagenmakers et al., 2010</xref>). Finally, Bayesian model averaging, stacking, and related methods can also help us address model uncertainty without the need to select a single model (e.g., <xref ref-type="bibr" rid="bid__13">Kaplan, 2021</xref>; <xref ref-type="bibr" rid="bid__46">Yao et al., 2018</xref>). We acknowledge that none of these solutions provides a fast and simple replacement for conditional information criteria that are automatically computed by some software.</p></sec>
<sec id="s6.2" sec-type="body"><title>Metric Selection</title>
    <p>Throughout the paper, we highlighted that a specific information criterion (marginal or otherwise) should be chosen based on the type of cross-validation that each criterion approximates. Other researchers sometimes focus on a criterion’s ability to select the data generating model. In our view, the model’s ability to select the true model is not necessarily relevant in situations where the true model does not exist. This topic was discussed at length in a special issue of <italic>Computational</italic> <italic>Brain &amp; Behavior</italic> by authors with diverse points of view (e.g., <xref ref-type="bibr" rid="bid__11">Gronau &amp; Wagenmakers, 2019</xref>; <xref ref-type="bibr" rid="bid__43">Vehtari et al., 2019</xref>). <xref ref-type="bibr" rid="bid__29">Piironen and Vehtari (2017)</xref> relatedly considered the use of different metrics depending on whether or not we expect the true model to be in the set of candidate models (see their Table 1). We also recall the earlier remarks of <xref ref-type="bibr" rid="bid__3">Browne (2000)</xref>: “Any attempt to use cross-validation to detect which of a set of models is correct disregards the fundamental intent of the approach. This is to find a model that yields as small an overall discrepancy as is possible given a specific sample size, not to find a model with no error of approximation. In other words, the intent is to find a model that yields a calibration that is worth further examination despite a possibly small sample, not to seek a correct model with an unknown calibration.” (p. 130).</p>
<p>More generally, as pointed out by <xref ref-type="bibr" rid="bid__27">Navarro (2019)</xref>, all the information criteria and other metrics mentioned in this paper emphasize quantitative measures of model performance at the neglect of whether any of the models help solve scientific problems. The overall goal of solving scientific problems should not be lost among model selection metrics.</p></sec>
<sec id="s6.3" sec-type="body"><title>Recommendations</title>
<p>We offer a variety of practical recommendations for researchers who estimate Bayesian models via MCMC and report Bayesian information criteria. First, verify that there are random model parameters (that is, parameters that would be called “random” in a frequentist context), which will often be called random effects, latent variables, or factors. If there are no random parameters, then the issues described in this paper are not applicable. Second, ensure that a large number of posterior samples are taken to reduce noise in the information criteria. We recommend starting with twice the number of samples that one would use for posterior means of individual parameters, as well as monitoring the effective sample size of the log-likelihood values that are used to compute information criteria. One may encounter a speed-accuracy tradeoff in this step, and conditional information criteria will have more noise than marginal.</p>
<p>Following estimation, researchers should verify the type of information criteria that their software supplies. To our knowledge, Mplus and <italic>blavaan</italic> are the only programs that supply marginal information criteria of psychometric models by default. Most other software will produce conditional information criteria because the computations are easier. The effective number of parameters associated with a criterion often provides a hint about what likelihood is being used: if the effective number of parameters is close to the frequentist count of model parameters, then a marginal likelihood is being used. Fourth, if one has conditional information criteria but does not have a good reason to prefer conditional information criteria, consider whether it is possible to compute marginal information criteria (with subsampling, if necessary), to compute a related cross-validation metric, or to estimate an encompassing model; see the Computational Issues section above. Finally, be transparent in reporting: provide details about the specific types of metrics that were reported, and be careful about using the metrics to select one model in light of the nuances described in this paper.</p></sec>
<sec id="s6.4" sec-type="body"><title>Conclusions</title>
<p>The popularity of Bayesian information criteria partly stems from the fact that they can be byproducts of MCMC estimation, allowing researchers to compare models without needing to consider post-estimation checks or computations. This simplicity becomes problematic as models increase in complexity, because many researchers do not expect the information criteria to require extra consideration or computations. In this paper, we highlighted some of these issues and made some recommendations for computing information criteria and related cross-validation metrics. The issues might be disappointing to some researchers because they imply that the information criteria require more work and attention than is desired. This could lead researchers to consider whether they are using information criteria because they are simple or because they are the best tools for their work. Because computation and effort are not infinite, researchers might sometimes find that their resources are better spent on posterior checks, assessment of qualitative differences between models, and encompassing models.</p></sec></sec>
<sec id="s7" sec-type="body"><title>Computational Details</title>
    <p>All results were obtained using the R system for statistical computing (<xref ref-type="bibr" rid="bid__31">R Core Team, 2025</xref>), Version 4.5.2, especially relying on packages <italic>blavaan</italic> (<xref ref-type="bibr" rid="bid__22">Merkle et al., 2021</xref>), <italic>loo</italic> (<xref ref-type="bibr" rid="bid__40">Vehtari et al., 2024</xref>), and <italic>rstan</italic> (<xref ref-type="bibr" rid="bid__37">Stan Development Team, 2025</xref>). The code for reproducing the results accompanies this submission (see <xref ref-type="bibr" rid="bid__21.25">Merkle, 2026a</xref>).</p></sec>
</body>
<back>
<ref-list><title>References</title>
<ref id="bid__1"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Asparouhov</surname>, <given-names>T.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Muthén</surname>, <given-names>B.</given-names></string-name> (<year>2020</year>). <article-title>Comparison of models for the analysis of intensive longitudinal data</article-title>. <source><italic>Structural Equation Modeling: A Multidisciplinary Journal</italic></source>, <volume>27</volume><issue>2</issue>, <fpage>275</fpage>–<lpage>297</lpage>. <pub-id pub-id-type="doi">10.1080/10705511.2019.1626733</pub-id></mixed-citation></ref>
    <ref id="bid__2"><mixed-citation publication-type="web"><string-name name-style="western"><surname>Bhattacjarjee</surname>, <given-names>S.</given-names></string-name> (<year>2016</year>). <italic>tmvnsim: Truncated multivariate normal simulation</italic> [R Package Version 1.0-2]. R Project for Statistical Computing. <ext-link ext-link-type="uri" xlink:href="https://CRAN.R-project.org/package=tmvnsim">https://CRAN.R-project.org/package=tmvnsim</ext-link></mixed-citation></ref>
    <ref id="bid__3"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Browne</surname>, <given-names>M. W.</given-names></string-name> (<year>2000</year>). <article-title>Cross-validation methods</article-title>. <source><italic>Journal of Mathematical Psychology</italic></source>, <volume>44</volume><issue>1</issue>, <fpage>108</fpage>–<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1006/jmps.1999.1279</pub-id></mixed-citation></ref>
    <ref id="bid__4"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Browne</surname>, <given-names>M. W.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Cudeck</surname>, <given-names>R.</given-names></string-name> (<year>1989</year>). <article-title>Single sample cross-validation indices for covariance structures</article-title>. <source><italic>Multivariate Behavioral Research</italic></source>, <volume>24</volume><issue>4</issue>, <fpage>445</fpage>–<lpage>455</lpage>.<pub-id pub-id-type="doi">10.1207/s15327906mbr2404_4</pub-id></mixed-citation></ref>
<ref id="bid__5"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Bürkner</surname>, <given-names>P.-C.</given-names></string-name> (<year>2017</year>). <article-title>brms: An R package for Bayesian multilevel models using Stan</article-title>. <source><italic>Journal of Statistical Software</italic></source>, <volume>80</volume>(<issue>1</issue>), <fpage>1</fpage>–<lpage>28</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v080.i01</pub-id></mixed-citation></ref>
    <ref id="bid__6"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Celeux</surname>, <given-names>G.</given-names></string-name>, <string-name name-style="western"><surname>Forbes</surname>, <given-names>F.</given-names></string-name>, <string-name name-style="western"><surname>Robert</surname>, <given-names>C. P.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Titterington</surname>, <given-names>D. M.</given-names></string-name>, (<year>2006</year>). <article-title>Deviance information criteria for missing data models</article-title>. <source><italic>Bayesian Analysis</italic></source>, <volume>1</volume>(<issue>4</issue>), <fpage>651</fpage>–<lpage>673</lpage>.<pub-id pub-id-type="doi">10.1214/06-BA122</pub-id></mixed-citation></ref>
<ref id="bid__7"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Cudeck</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Browne</surname>, <given-names>M. W.</given-names></string-name> (<year>1983</year>). <article-title>Cross-validation of covariance structures</article-title>. <source><italic>Multivariate Behavioral Research</italic></source>, <volume>18</volume>(<issue>2</issue>), <fpage>147</fpage>–<lpage>167</lpage>. <pub-id pub-id-type="doi">10.1207/s15327906mbr1802_2</pub-id></mixed-citation></ref>
<ref id="bid__8"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Du</surname>, <given-names>H.</given-names></string-name>, <string-name name-style="western"><surname>Keller</surname>, <given-names>B.</given-names></string-name>, <string-name name-style="western"><surname>Alacam</surname>, <given-names>E.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Enders</surname>, <given-names>C.</given-names></string-name> (<year>2023</year>). <article-title>Comparing DIC and WAIC for multilevel models with missing data</article-title>. <source><italic>Behavior Research Methods</italic></source>, <volume>56</volume>(<issue>4</issue>), <fpage>2731</fpage>–<lpage>2750</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-023-02231-0</pub-id></mixed-citation></ref>
<ref id="bid__9"><mixed-citation publication-type="book"><string-name name-style="western"><surname>du Toit</surname>, <given-names>S. H. C.</given-names></string-name>, &amp; <string-name name-style="western"><surname>du Toit</surname>, <given-names>M.</given-names></string-name> (<year>2008</year>). <chapter-title>Multilevel structural equation modeling</chapter-title>. In J. De Leeuw &amp; E. Meijer (Eds.), <italic>Handbook of multilevel analysis</italic> (pp. 435–478). <pub-id pub-id-type="doi">10.1007/978-0-387-73186-5_12</pub-id></mixed-citation></ref>
    <ref id="bid__10"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Graves</surname>, <given-names>B.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name> (<year>2022</year>). <article-title>A note on identification constraints and information criteria in Bayesian latent variable models</article-title>. <source><italic>Behavior Research Methods</italic></source>, <volume>54</volume>, <fpage>795</fpage>–<lpage>804</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-021-01649-8</pub-id></mixed-citation></ref>
    <ref id="bid__11"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Gronau</surname>, <given-names>Q. F.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Wagenmakers</surname>, <given-names>E.-J.</given-names></string-name> (<year>2019</year>). <article-title>Limitations of Bayesian leave-one-out cross-validation for model selection</article-title>. <source><italic>Computational Brain &amp; Behavior</italic></source>, <volume>2</volume>(<issue>1</issue>), <fpage>1</fpage>–<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1007/s42113-018-0011-7</pub-id></mixed-citation></ref>
<ref id="bid__12"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Hajivassiliou</surname>, <given-names>V. A.</given-names></string-name>, &amp; <string-name name-style="western"><surname>McFadden</surname>, <given-names>D. L.</given-names></string-name> (<year>1998</year>). <article-title>The method of simulated scores for the estimation of LDV models</article-title>. <source><italic>Econometrica</italic></source>, <volume>66</volume>(<issue>4</issue>), <fpage>863</fpage>–<lpage>896</lpage>. doi:<pub-id pub-id-type="doi">10.2307/2999576</pub-id></mixed-citation></ref>
    <ref id="bid__13"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Kaplan</surname>, <given-names>D.</given-names></string-name> (<year>2021</year>). <article-title>On the quantification of model uncertainty: A Bayesian perspective</article-title>. <source><italic>Psychometrika</italic></source>, <volume>86</volume><issue>1</issue>, <fpage>215</fpage>–<lpage>238</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-021-09754-5</pub-id></mixed-citation></ref>
    <ref id="bid__14"><mixed-citation publication-type="web"><string-name name-style="western"><surname>Keller</surname>, <given-names>B. T.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Enders</surname>, <given-names>C. K.</given-names></string-name> (<year>2023</year>). <source><italic>Blimp user’s guide</italic></source> (3<sup>rd</sup> ed.). <ext-link ext-link-type="uri" xlink:href="https://www.appliedmissingdata.com/blimp">https://www.appliedmissingdata.com/blimp</ext-link>.</mixed-citation></ref>
    <ref id="bid__15"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Lee</surname>, <given-names>S.-Y.</given-names></string-name>, <string-name name-style="western"><surname>Song</surname>, <given-names>X.-Y.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Tang</surname>, <given-names>N.-S.</given-names></string-name> (<year>2007</year>). <article-title>Bayesian methods for analyzing structural equation models with covariates, interaction, and quadratic latent variables</article-title>. <source><italic>Structural Equation Modeling</italic></source>, <volume>14</volume><issue>3</issue>, <fpage>404</fpage>–<lpage>434</lpage>. <pub-id pub-id-type="doi">10.1080/10705510701301511</pub-id></mixed-citation></ref>
    <ref id="bid__16"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Li</surname>, <given-names>L.</given-names></string-name>, <string-name name-style="western"><surname>Qui</surname>, <given-names>S.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Feng</surname>, <given-names>C. X.</given-names></string-name> (<year>2016</year>). <article-title>Approximating cross-validatory predictive evaluation in Bayesian latent variable models with integrated IS and WAIC</article-title>. <source><italic>Statistics and Computing</italic></source>, <volume>26</volume><issue>4</issue>, <fpage>881</fpage>–<lpage>897</lpage>. <pub-id pub-id-type="doi">10.1007/s11222-015-9577-2</pub-id></mixed-citation></ref>
<ref id="bid__17"><mixed-citation publication-type="web"><string-name name-style="western"><surname>Luo</surname>, <given-names>X.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Dong</surname>, <given-names>J.</given-names></string-name> (<year>2025</year>). <italic>bleval: Bayesian evaluation for latent variable models</italic> [R package Version 0.0.0.9000]. GitHub. <ext-link ext-link-type="uri" xlink:href="https://github.com/luoxh3/bleval">https://github.com/luoxh3/bleval</ext-link></mixed-citation></ref>
<ref id="bid__18"><mixed-citation publication-type="author"><string-name name-style="western"><surname>Luo</surname>, <given-names>X.</given-names></string-name>, <string-name name-style="western"><surname>Dong</surname>, <given-names>J.</given-names></string-name>, <string-name name-style="western"><surname>Liu</surname>, <given-names>H.</given-names></string-name>, <string-name name-style="western"><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name> (<year>2026</year>). <italic>Bayesian evaluation for latent variable models: A tutorial on computing information criteria and Bayes factors with the R package</italic> bleval [Manuscript in preparation].</mixed-citation></ref>
<ref id="bid__19"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Magnusson</surname>, <given-names>M.</given-names></string-name>, <string-name name-style="western"><surname>Andersen</surname>, <given-names>M. R.</given-names></string-name>, <string-name name-style="western"><surname>Jonasson</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name> (<year>2020</year>). <italic>Leave-one-out cross-validation for Bayesian model comparison in large data</italic> [arXiv 2001.00980]. arXiv. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2001.00980">https://arxiv.org/abs/2001.00980</ext-link></mixed-citation></ref>
<ref id="bid__20"><mixed-citation publication-type="web"><string-name name-style="western"><surname>Magnusson</surname>, <given-names>M.</given-names></string-name>, <string-name name-style="western"><surname>Bürkner</surname>, <given-names>P.</given-names></string-name>, <string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Gabry</surname>, <given-names>J.</given-names></string-name> (2025, December 23). <article-title>Using leave-one-out cross-validation for large data</article-title>. <ext-link ext-link-type="uri" xlink:href="https://mc-stan.org/loo/articles/loo2-large-data.html">https://mc-stan.org/loo/articles/loo2-large-data.html</ext-link></mixed-citation></ref>
<ref id="bid__21"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>McDonald</surname>, <given-names>R. P.</given-names></string-name> (<year>1993</year>). <article-title>A general model for two-level data with responses missing at random</article-title>. <source><italic>Psychometrika</italic></source>, <volume>58</volume>(<issue>4</issue>), <fpage>575</fpage>–<lpage>585</lpage>. <pub-id pub-id-type="doi">10.1007/bf02294828</pub-id></mixed-citation></ref>

    <ref id="bid__21.25"><mixed-citation publication-type="data"><string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name> (<year>2026a</year>). <data-title><italic>Supplementary Materials to</italic> “Nuances of information criteria for Bayesian psychometric models”</data-title> <comment>[R codes]</comment>. PsychOpen GOLD. <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.23668/psycharchives.22424">http://dx.doi.org/10.23668/psycharchives.22424</ext-link></mixed-citation></ref>
    <ref id="bid__21.75"><mixed-citation publication-type="data"><string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name> (<year>2026b</year>). <data-title><italic>Supplementary Materials to</italic> “Nuances of information criteria for Bayesian psychometric models”</data-title> <comment>[Full study data]</comment>. PsychOpen GOLD. <ext-link ext-link-type="uri" xlink:href="http://dx.doi.org/10.23668/psycharchives.22423">http://dx.doi.org/10.23668/psycharchives.22423</ext-link></mixed-citation></ref>
    
<ref id="bid__22"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name>, <string-name name-style="western"><surname>Fitzsimmons</surname>, <given-names>E.</given-names></string-name>, <string-name name-style="western"><surname>Uanhoro</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Goodrich</surname>, <given-names>B.</given-names></string-name> (<year>2021</year>). <article-title>Efficient Bayesian structural equation modeling in Stan</article-title>. <source><italic>Journal of Statistical Software</italic></source>, <volume>100</volume>(<issue>6</issue>), <fpage>1</fpage>–<lpage>22</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v100.i06</pub-id></mixed-citation></ref>
    <ref id="bid__23"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name>, <string-name name-style="western"><surname>Furr</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Rabe-Hesketh</surname>, <given-names>S.</given-names></string-name> (<year>2019</year>). <article-title>Bayesian comparison of latent variable models: Conditional versus marginal likelihoods</article-title>. <source><italic>Psychometrika</italic></source>, <volume>84</volume><issue>3</issue>, <fpage>802</fpage>–<lpage>829</lpage>. <pub-id pub-id-type="doi">10.1007/s11336-019-09679-0</pub-id></mixed-citation></ref>
    <ref id="bid__24"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Rosseel</surname>, <given-names>Y.</given-names></string-name> (<year>2018</year>). <article-title>blavaan: Bayesian structural equation models via parameter expansion</article-title>. <source><italic>Journal of Statistical Software</italic></source>, <volume>85</volume>(<issue>4</issue>), <fpage>1</fpage>–<lpage>30</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v085.i04</pub-id></mixed-citation></ref>
<ref id="bid__25"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name>, <string-name name-style="western"><surname>You</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Preacher</surname>, <given-names>K. J.</given-names></string-name> (<year>2016</year>). <article-title>Testing non-nested structural equation models</article-title>. <source><italic>Psychological Methods</italic></source>, <volume>21</volume><issue>2</issue>, <fpage>151</fpage>–<lpage>163</lpage>. <pub-id pub-id-type="doi">10.1037/met0000038</pub-id></mixed-citation></ref>
    <ref id="bid__26"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Millar</surname>, <given-names>R. B.</given-names></string-name> (<year>2018</year>). <article-title>Conditional vs. marginal estimation of predictive loss of hierarchical models using WAIC and cross-validation</article-title>. <source><italic>Statistics and Computing</italic></source>, <volume>28</volume><issue>2</issue>, <fpage>375</fpage>–<lpage>385</lpage>. <pub-id pub-id-type="doi">10.1007/s11222-017-9736-8</pub-id></mixed-citation></ref>
    <ref id="bid__27"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Navarro</surname>, <given-names>D. J.</given-names></string-name> (<year>2019</year>). <article-title>Between the devil and the deep blue sea: Tensions between scientific judgement and statistical model selection</article-title>. <source><italic>Computational Brain &amp; Behavior</italic></source>, <volume>2</volume>, <fpage>28</fpage>–<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1007/s42113-018-0019-z</pub-id></mixed-citation></ref>
    <ref id="bid__28"><mixed-citation publication-type="web"><collab>Open Source Psychometrics Project</collab> (<year>2016</year>). <italic>Data from</italic> “The Nerdy Personality Attributes Scale” [Dataset]. <ext-link ext-link-type="uri" xlink:href="https://openpsychometrics.org/%5C_rawdata/">https://openpsychometrics.org/%5C_rawdata/</ext-link></mixed-citation></ref>
    <ref id="bid__29"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Piironen</surname>, <given-names>J.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name> (<year>2017</year>). <article-title>Comparison of Bayesian predictive methods for model selection</article-title>. <source><italic>Statistics and Computing</italic></source>, <volume>27</volume>, <fpage>711</fpage>–<lpage>735</lpage>. <pub-id pub-id-type="doi">10.1007/s11222-016-9649-y</pub-id></mixed-citation></ref>
    <ref id="bid__30"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Plummer</surname>, <given-names>M.</given-names></string-name> (<year>2008</year>). <article-title>Penalized loss functions for Bayesian model comparison</article-title>. <source><italic>Biostatistics</italic></source>, <volume>9</volume>(<issue>3</issue>), <fpage>523</fpage>–<lpage>539</lpage>. <pub-id pub-id-type="doi">10.1093/biostatistics/kxm049</pub-id></mixed-citation></ref>
<ref id="bid__31"><mixed-citation publication-type="web"><collab>R Core Team</collab> (<year>2025</year>). <source><italic>R: A language and environment for statistical computing</italic></source>. R Foundation for Statistical Computing. <ext-link ext-link-type="uri" xlink:href="https://www.R-project.org/">https://www.R-project.org/</ext-link></mixed-citation></ref>
    <ref id="bid__32"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Rabe-Hesketh</surname>, <given-names>S.</given-names></string-name>, <string-name name-style="western"><surname>Skrondal</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Pickles</surname>, <given-names>A.</given-names></string-name> (<year>2005</year>). <article-title>Maximum likelihood estimation of limited and discrete dependent variable models with nested random effects</article-title>. <source><italic>Journal of Econometrics</italic></source>, <volume>128</volume>(<issue>2</issue>), <fpage>301</fpage>–<lpage>323</lpage>. <pub-id pub-id-type="doi">10.1016/j.jeconom.2004.08.017</pub-id></mixed-citation></ref>
<ref id="bid__33"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Rosseel</surname>, <given-names>Y.</given-names></string-name> (<year>2021</year>). <article-title>Evaluating the observed log-likelihood function in two-level structural equation modeling with missing data: From formulas to R code</article-title>. <source><italic>Psych</italic></source>, <volume>3</volume>(<issue>2</issue>), <fpage>197</fpage>–<lpage>232</lpage>. <pub-id pub-id-type="doi">10.3390/psych3020017</pub-id></mixed-citation></ref>
    <ref id="bid__34"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Schneider</surname>, <given-names>L.</given-names></string-name>, <string-name name-style="western"><surname>Chalmers</surname>, <given-names>R. P.</given-names></string-name>, <string-name name-style="western"><surname>Debelak</surname>, <given-names>R.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Merkle</surname>, <given-names>E. C.</given-names></string-name> (<year>2020</year>). <article-title>Model selection of nested and non-nested item response models using Vuong tests</article-title>. <source><italic>Multivariate Behavioral Research</italic></source>, <volume>55</volume><issue>5</issue>, <fpage>664</fpage>–<lpage>684</lpage>. <pub-id pub-id-type="doi">10.1080/00273171.2019.1664280</pub-id></mixed-citation></ref>
<ref id="bid__35"><mixed-citation publication-type="book"><string-name name-style="western"><surname>Skrondal</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Rabe-Hesketh</surname>, <given-names>S.</given-names></string-name> (<year>2004</year>). <source><italic>Generalized latent variable modeling: multilevel, longitudinal, and structural equation modeling</italic></source>. <publisher-name>Chapman &amp; Hall</publisher-name>.</mixed-citation></ref>
    <ref id="bid__36"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Spiegelhalter</surname>, <given-names>D. J.</given-names></string-name>, <string-name name-style="western"><surname>Best</surname>, <given-names>N. G.</given-names></string-name>, <string-name name-style="western"><surname>Carlin</surname>, <given-names>B. P.</given-names></string-name>, &amp; <string-name name-style="western"><surname>van der Linde</surname>, <given-names>A.</given-names></string-name> (<year>2002</year>). <article-title>Bayesian measures of model complexity and fit</article-title>. <source>Journal of the Royal Statistical Society; Series B (Statistical Methodology)</source>, <volume>64</volume><issue>4</issue>, <fpage>583</fpage>–<lpage>639</lpage>. <pub-id pub-id-type="doi">10.1111/1467-9868.00353</pub-id></mixed-citation></ref>
<ref id="bid__37"><mixed-citation publication-type="web"><collab>Stan Development Team</collab> (<year>2025</year>). <source>RStan: The R interface to Stan</source> [R package Version 2.32.7]. Stan Project. <ext-link ext-link-type="uri" xlink:href="https://mc-stan.org/">https://mc-stan.org/</ext-link></mixed-citation></ref>
    <ref id="bid__38"><mixed-citation publication-type="other"><string-name name-style="western"><surname>Tanner</surname>, <given-names>M. A.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Wong</surname>, <given-names>W. H.</given-names></string-name> (<year>1987</year>). <article-title>The calculation of posterior distributions by data augmentation</article-title>. <source> Journal of the American Statistical Association</source>, <volume>82</volume><issue>398</issue>, <fpage>528</fpage>–<lpage>540</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1987.10478458</pub-id></mixed-citation></ref>
    <ref id="bid__39"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Timonen</surname>, <given-names>J.</given-names></string-name>, <string-name name-style="western"><surname>Siccha</surname>, <given-names>N.</given-names></string-name>, <string-name name-style="western"><surname>Bales</surname>, <given-names>B.</given-names></string-name>, <string-name name-style="western"><surname>Lähdesmäki</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name> (<year>2023</year>). <article-title>An importance sampling approach for reliable and efficient inference in Bayesian ordinary differential equation models</article-title>. <source><italic>Stat</italic></source>, <volume>12</volume>(<issue>1</issue>), <elocation-id>e614</elocation-id>. <pub-id pub-id-type="doi">10.1002/sta4.614</pub-id></mixed-citation></ref>
<ref id="bid__40"><mixed-citation publication-type="web"><string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name>, <string-name name-style="western"><surname>Gabry</surname>, <given-names>J.</given-names></string-name>, <string-name name-style="western"><surname>Magnusson</surname>, <given-names>M.</given-names></string-name>, <string-name name-style="western"><surname>Yao</surname>, <given-names>Y.</given-names></string-name>, <string-name name-style="western"><surname>Bürkner</surname>, <given-names>P.-C.</given-names></string-name>, <string-name name-style="western"><surname>Paananen</surname>, <given-names>T.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Gelman</surname>, <given-names>A.</given-names></string-name> (<year>2024</year>). <italic>loo: Efficient leave-one-out cross-validation and WAIC for Bayesian models</italic> [R package Version 2.8.0]. Stan Project. <ext-link ext-link-type="uri" xlink:href="https://mc-stan.org/loo/">https://mc-stan.org/loo/</ext-link></mixed-citation></ref>
    <ref id="bid__41"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name>, <string-name name-style="western"><surname>Gelman</surname>, <given-names>A.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Gabry</surname>, <given-names>J.</given-names></string-name> (<year>2017</year>). <article-title>Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC</article-title>. <source><italic>Statistics and Computing</italic></source>, <volume>27</volume><issue>5</issue>, <fpage>1413</fpage>–<lpage>1432</lpage>. <pub-id pub-id-type="doi">10.1007/s11222-016-9696-4</pub-id></mixed-citation></ref>
<ref id="bid__42"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name>, <string-name name-style="western"><surname>Mononen</surname>, <given-names>T.</given-names></string-name>, <string-name name-style="western"><surname>Tolvanen</surname>, <given-names>V.</given-names></string-name>, <string-name name-style="western"><surname>Sivula</surname>, <given-names>T.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Winther</surname>, <given-names>O.</given-names></string-name> (<year>2016</year>). <article-title>Bayesian leave-one-out cross-validation approximations for Gaussian latent variable models</article-title>. <source><italic>Journal of Machine Learning Research</italic></source>, <volume>17</volume><issue>103</issue>, <fpage>1</fpage>–<lpage>38</lpage>.</mixed-citation></ref>
    <ref id="bid__43"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name>, <string-name name-style="western"><surname>Simpson</surname>, <given-names>D. P.</given-names></string-name>, <string-name name-style="western"><surname>Yao</surname>, <given-names>Y.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Gelman</surname>, <given-names>A.</given-names></string-name> (<year>2019</year>). <article-title>Limitations of “Limitations of Bayesian leave-one-out cross-validation for model selection”</article-title>. <source><italic>Computational Brain &amp; Behavior</italic></source>, <volume>2</volume>(<issue>1</issue>), <fpage>22</fpage>–<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1007/s42113-018-0020-6</pub-id></mixed-citation></ref>
    <ref id="bid__44"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Wagenmakers</surname>, <given-names>E.-J.</given-names></string-name>, <string-name name-style="western"><surname>Lodewyckx</surname>, <given-names>T.</given-names></string-name>, <string-name name-style="western"><surname>Kuriyal</surname>, <given-names>H.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Grasman</surname>, <given-names>R.</given-names></string-name> (<year>2010</year>). <article-title>Bayesian hypothesis testing for psychologists: A tutorial on the Savage-Dickey method</article-title>. <source><italic>Cognitive Psychology</italic></source>, <volume>60</volume><issue>3</issue>, <fpage>158</fpage>–<lpage>189</lpage>. <pub-id pub-id-type="doi">10.1016/j.cogpsych.2009.12.001</pub-id></mixed-citation></ref>
    <ref id="bid__45"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Watanabe</surname>, <given-names>S.</given-names></string-name> (<year>2010</year>). <article-title>Asymptotic equivalence of Bayes cross validation and widely applicable information criterion in singular learning theory</article-title>. <source><italic>Journal of Machine Learning Research</italic></source>, <volume>11</volume>, <fpage>3571</fpage>–<lpage>3594</lpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1004.2316</pub-id></mixed-citation></ref>
<ref id="bid__46"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Yao</surname>, <given-names>Y.</given-names></string-name>, <string-name name-style="western"><surname>Vehtari</surname>, <given-names>A.</given-names></string-name>, <string-name name-style="western"><surname>Simpson</surname>, <given-names>D.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Gelman</surname>, <given-names>A.</given-names></string-name> (<year>2018</year>). <article-title>Using stacking to average Bayesian predictive distributions (with discussion)</article-title>. <source><italic>Bayesian Analysis</italic></source>, <volume>13</volume><issue>3</issue>, <fpage>917</fpage>–<lpage>1007</lpage>. <pub-id pub-id-type="doi">10.1214/17-BA1091</pub-id></mixed-citation></ref>
<ref id="bid__47"><mixed-citation publication-type="journal"><string-name name-style="western"><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name name-style="western"><surname>Tao</surname>, <given-names>J.</given-names></string-name>, <string-name name-style="western"><surname>Wang</surname>, <given-names>C.</given-names></string-name>, &amp; <string-name name-style="western"><surname>Shi</surname>, <given-names>N.-Z.</given-names></string-name> (<year>2019</year>). <article-title>Bayesian model selection methods for multilevel IRT models: A comparison of five DIC-based indices</article-title>. <source><italic>Journal of Educational Measurement</italic></source>, <volume>56</volume>(<issue>1</issue>), <fpage>3</fpage>–<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1111/jedm.12197</pub-id></mixed-citation></ref>
</ref-list>
<app-group>
<app id="app"><title>Appendix</title>
<sec id="s9"><title>Appendix A: Clusterwise Marginal Likelihood Contributions in Two-Level SEM With Random Intercepts</title>
<p>In this appendix, we provide technical details about how to obtain cluster-level contributions to the marginal likelihood of a two-level SEM with random intercepts. Many relevant derivations are presented by <xref ref-type="bibr" rid="bid__33">Rosseel (2021)</xref>, who builds on work by <xref ref-type="bibr" rid="bid__9">du Toit and du Toit (2008)</xref> and <xref ref-type="bibr" rid="bid__21">McDonald (1993)</xref>. For simplicity, we exclude observed variables that only occur at Level 2 (e.g., school-level observed variables). For intuitive discussion, we generically refer to Level 1 units as “students” and Level 2 clusters as “schools.”</p>
<p>The model is defined by Equations <xref ref-type="disp-formula" rid="xd12">(12)</xref> and <xref ref-type="disp-formula" rid="xd15">(15)</xref>. The model from those equations includes the latent variables at both levels, and we wish to obtain a likelihood that is marginalized over those latent variables. But in marginalizing over the latent variables, we must account for correlations between students who are in the same school. Because our model involves multivariate normal distributions, we can obtain analytic expressions for the marginal likelihood.</p>
<p>An initial step towards these analytic expressions involves obtaining the mean and covariance of <italic><bold>y</bold><sub>ik</sub></italic>, i.e., the mean and covariance of the observed variables from a particular student. These are similar to expressions from one-level SEM and are given by</p>
<p><disp-formula id="xd21"><label>21</label><mml:math id="xa-21" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">μ</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">ν</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="bold">Λ</mml:mi><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold-italic">α</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd22"><label>22</label><mml:math id="xa-22" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where the covariance matrix <bold>Σ</bold><italic><sub>y</sub></italic> is familiarly decomposed into within and between components. Some treatments of multilevel SEM with random intercepts define the model as this decomposition. This may be useful when the mean vector is not of interest, but it is also not the most intuitive model definition. Regardless, the between and within components mirror one another, involving parameter matrices at the student and school levels, respectively:</p>
<p><disp-formula id="xd23"><label>23</label><mml:math id="xa-23" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi mathvariant="bold-italic">w</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="bold">Λ</mml:mi><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mi mathvariant="bold">Ψ</mml:mi><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:msup><mml:mn>1</mml:mn><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mo>′</mml:mo></mml:msup><mml:mo>+</mml:mo><mml:mi mathvariant="bold">Θ</mml:mi></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd24"><label>24</label><mml:math id="xa-24" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi mathvariant="bold-italic">b</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold">Ψ</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">I</mml:mi><mml:mo>−</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:msup><mml:mn>1</mml:mn><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold">Λ</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>′</mml:mo></mml:msup></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi mathvariant="bold">Θ</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula></p>
<p>Next, let <italic>n<sub>k</sub></italic> be the number of individuals in school <italic>k</italic>. Then our marginal likelihood for all observed variables in school <italic>k</italic> is multivariate normal with mean</p>
<p><disp-formula id="xd25"><label>25</label><mml:math id="jo-21" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mtable columnspacing="1em" rowspacing="4pt"><mml:mtr><mml:mtd><mml:mrow><mml:msubsup><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mi>′</mml:mi></mml:msubsup><mml:mo>⊗</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi><mml:mi>′</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo>)</mml:mo></mml:mrow><mml:mi>′</mml:mi></mml:msup></mml:math></disp-formula></p>
<p>and covariance matrix</p>
<p><disp-formula id="xd26"><label>26</label><mml:math id="jo-22" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">V</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mtable columnspacing="1em" rowspacing="4pt"><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>⁢</mml:mo><mml:msubsup><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mi>′</mml:mi></mml:msubsup></mml:mrow><mml:mo>⊗</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">I</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>⊗</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>where <inline-formula id="asdfg1"><mml:math id="aju"><mml:msub><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> is a column vector of <italic>n<sub>k</sub></italic> 1s. This leads us to express school <italic>k</italic>’s marginal log-likelihood as</p>
<p><disp-formula id="xd27"><label>27</label><mml:math id="jo-23" display="block"><mml:mrow><mml:mrow><mml:msub><mml:mi>ℓ</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">θ</mml:mi></mml:mrow><mml:mo>∣</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>−</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>⁢</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:mrow><mml:mi>log</mml:mi><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mo>⁢</mml:mo><mml:mi>π</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>log</mml:mi><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">V</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>′</mml:mi></mml:msup><mml:mo>⁢</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">V</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>where <italic>P<sub>k</sub></italic> = <italic>n<sub>k</sub>p</italic>.</p>
    <p>While the above expression is all we need to evaluate the marginal log-likelihood contribution of each school, it will usually be very slow to evaluate. It involves both the inverse and determinant of <italic><bold>V</bold><sub>k</sub></italic>, with the dimension of this matrix equaling the total number of observed variables in school <italic>k</italic> (<italic>P<sub>k</sub></italic>). But the structure of <italic><bold>V</bold><sub>k</sub></italic> allows for a variety of simplifications (see <xref ref-type="bibr" rid="bid__33">Rosseel, 2021</xref>). In particular, the log-determinant can be written as</p>
<p><disp-formula id="xd28"><label>28</label><mml:math id="jo-24" display="block"><mml:mrow><mml:mrow><mml:mi>log</mml:mi><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">V</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>⁢</mml:mo><mml:mrow><mml:mi>log</mml:mi><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi></mml:msub><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>log</mml:mi><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>which involves determinants of much smaller matrices. Relatedly, we have</p>
<p><disp-formula id="xd29"><label>29</label><mml:math id="jo-25" display="block"><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>′</mml:mi></mml:msup><mml:mo>⁢</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">V</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mspace width="0.5em"/></mml:mrow></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mrow><mml:mtext>tr</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">Y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>⊗</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi><mml:mi>′</mml:mi></mml:msubsup></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>⁢</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>⁢</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">Y</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>⊗</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi><mml:mi>′</mml:mi></mml:msubsup></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>′</mml:mi></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>⁢</mml:mo><mml:mrow><mml:mspace width="0.5em"/><mml:mo>−</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mstyle indentshift="2em"><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>′</mml:mi></mml:msup><mml:mo>⁢</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>⁢</mml:mo><mml:mrow><mml:mspace width="0.5em"/><mml:mo>+</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mstyle indentshift="2em"><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mi>′</mml:mi></mml:msup><mml:mo>⁢</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>⁢</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>−</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>⁢</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mover><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo accent="true">―</mml:mo></mml:mover><mml:mi>k</mml:mi></mml:msub><mml:mo>−</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">μ</mml:mi></mml:mrow><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>where <italic><bold>Y</bold><sub>k</sub></italic> is an <italic>n<sub>k</sub></italic> × <italic>p</italic> matrix of observations in Cluster <italic>k</italic>. This expression involves inverses of <bold>Σ</bold><italic><sub>w</sub></italic> and <bold>Σ</bold><italic><sub>b</sub></italic>, both of which are of much smaller dimension than <italic><bold>V</bold><sub>k</sub></italic>. (Note that the <italic><bold>Y</bold><sub>k</sub></italic> matrix should not be confused with <italic><bold>y</bold><sub>k</sub></italic>, with the latter being a vector of length <italic>P<sub>k</sub></italic> containing all observations in Cluster <italic>k</italic>.)</p>
<p>As discussed by <xref ref-type="bibr" rid="bid__33">Rosseel (2021)</xref>, additional simplifications are available when we wish to compute the overall log-likelihood across schools. These additional simplifications are especially worthwhile when <italic>n<sub>k</sub></italic> is equal for most or all <italic>k</italic>. The expressions can also be modified to handle Level-2 observed variables and missing data. In addition to being used for information criterion computations, the expressions can be used for fast and efficient MCMC estimation as demonstrated by recent versions of <italic>blavaan</italic>.</p></sec>
<sec id="s10"><title>Appendix B: “Within-Only” Variables in Two-Level SEM</title>
<p>In this section, we illustrate how “within-only” variables in two-level SEM can lead to unexpected behavior of PSIS-LOO. As an example of a within-only variable, consider a situation where students complete three tests and also report the amount of anxiety that they felt during the tests. We might attribute test anxiety exclusively to individual students, as there is no reason to believe that anxiety is a property of a school. In this case, the observed anxiety variable still appears in Equation <xref ref-type="disp-formula" rid="xd15">(15)</xref>, but the corresponding row of <bold>Λ</bold><sup>(<italic>c</italic>)</sup> and the corresponding entry of <inline-formula><mml:math id="jo-26" display="inline"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold-italic">ϵ</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> are fixed to 0. When this “within-only” observed variable actually exhibits meaningful across-school variability in the observed data, the effective number of parameters in WAIC and PSIS-LOO are again inflated.</p>
<sec><title>Method</title>
<p>We generated data from a model with 4 observed variables and 200 clusters of size 50. The data-generating model can be written as:</p>
<p><disp-formula id="xd30"><label>30</label><mml:math id="jo-27" display="block"><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>⁢</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>∣</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">ν</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>N</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>ν</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.5em"/><mml:mspace width="0.5em"/><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>50</mml:mn><mml:mo>;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>200</mml:mn></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><disp-formula id="xd31"><label>31</label><mml:math id="jo-28" display="block"><mml:mtable><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi mathvariant="bold-italic">ν</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mrow><mml:mrow><mml:mrow/><mml:mo>∼</mml:mo><mml:mrow><mml:mtext>N</mml:mtext><mml:mo>⁡</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mn mathvariant="bold">0</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mspace width="0.5em"/><mml:mspace width="0.5em"/><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mo>…</mml:mo><mml:mo>,</mml:mo><mml:mn>200</mml:mn></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
    <p><disp-formula id="xd32"><label>32</label><mml:math id="xa-32" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="bold">Σ</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>0.3</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mn>1.5</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="xd33"><label>33</label><mml:math id="xa-33" display="block"><mml:mtable><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi mathvariant="bold">Σ</mml:mi></mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mstyle indentshift="2em"><mml:mi/><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mtable columnspacing="0.45em" rowspacing="4pt"><mml:mtr><mml:mtd><mml:mn>0.6</mml:mn></mml:mtd><mml:mtd><mml:mn>0.2</mml:mn></mml:mtd><mml:mtd><mml:mn>0.2</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0.2</mml:mn></mml:mtd><mml:mtd><mml:mn>0.6</mml:mn></mml:mtd><mml:mtd><mml:mn>0.2</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0.2</mml:mn></mml:mtd><mml:mtd><mml:mn>0.2</mml:mn></mml:mtd><mml:mtd><mml:mn>0.6</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:msub><mml:mi>σ</mml:mi><mml:mrow><mml:mn>44</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>)</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
    <p>where we manipulated the value of σ<sub>44</sub> in <bold>Σ</bold><italic><sub>b</sub></italic>. The σ<sub>44</sub> parameter assumed values from 0 to 4 in steps of 0.5, where a new dataset was generated for each value of σ<sub>44</sub>. We then fit a model where the fourth observed variable, <italic>y</italic><sub>4</sub>, is treated as “within only.” In <italic>(b)lavaan</italic>, this “within only” model could be specified as</p><preformat position="anchor" orientation="portrait" xml:space="preserve">
model &lt;-' 
    level: within
        fw =~ y1 + y2 + y3 + y4
    level: between
        fb =~ y1 + y2 + y3 
'
</preformat>
<p>Prior distributions were the same as those used in Example 2. For each of three chains, we used <italic>blavaan</italic> to obtain 2,000 samples following 500 burnin samples.</p></sec>
<sec><title>Results</title>
    <p>Results are shown in <xref ref-type="fig" rid="fB1">Figure B1</xref>. The true “between” variance of <italic>y</italic><sub>4</sub> is on the x-axis, effective number of parameter estimates are on the y-axis, and panels represent different information criteria. The horizontal line in each panel is the frequentist parameter count, which is 18 for this model. Similar to Example 2, we see that the DIC effective number of parameters is nearly flat and stays near the frequentist parameter count of 18. But under PSIS-LOO, the model appears to gain parameters as the between variability increases.</p><fig id="fB1" position="anchor" fig-type="figure"  orientation="portrait"><label>Figure B1</label><caption><title>Magnitude of Between Variability in “Within-Only” Variable (X-Axis) vs. Estimated Effective Number of Parameters (Y-Axis) of DIC and PSIS-LOO</title></caption><graphic xlink:href="meth.20361-fb1.png" position="anchor" orientation="portrait"/></fig></sec></sec>
</app>
</app-group>
        <sec sec-type="data-availability" id="das"><title>Data Availability</title>
            <p>The code (see <xref ref-type="bibr" rid="bid__21.25">Merkle, 2026a</xref>) and data (see <xref ref-type="bibr" rid="bid__21.75">Merkle, 2026b</xref>) for the paper are openly available in the Supplementary Materials.</p>
        </sec>
        
        
        
        <sec sec-type="supplementary-material" id="sp1"><title>Supplementary Materials</title>
            <table-wrap position="anchor" content-type="supplementary-materials">
                <table frame="void" style="background-#f3f3f3 nobreak">
                    <col width="60%" align="left"/>
                    <col width="40%" align="left"/>
                    <thead>
                        <tr>
                            <th>Type of supplementary material</th>
                            <th>Availability/Access</th>
                        </tr></thead>
                    <tbody>
                        <tr>
                            <th colspan="2">Data</th>						
                        </tr>
                        <tr><td>Merkle_2026_Nuances_SUPPL_data</td>
                            <td><xref ref-type="bibr" rid="bid__21.75">Merkle (2026b)</xref></td>
                        </tr>
                        <tr style="grey-border-top-dashed">
                            <th colspan="2">Code</th>
                        </tr>
                        <tr>
                            <td>Merkle_2026_Nuances_SUPPL_code.R</td>
                            <td><xref ref-type="bibr" rid="bid__21.25">Merkle (2026a)</xref></td>
                        </tr>
                        <tr style="grey-border-top-dashed">
                            <th colspan="2">Material</th>
                        </tr>
                        <tr>
                            <td>No supplementary material provided</td>
                            <td>&mdash;</td>
                        </tr>
                        <tr style="grey-border-top-dashed">
                            <th colspan="2">Study/Analysis preregistration</th>
                        </tr>	
                        <tr>
                            <td>No preregistration</td>
                            <td>&mdash;</td>
                        </tr>
                        <tr style="grey-border-top-dashed">
                            <th colspan="2">Other</th>
                        </tr>	
                        <tr>
                            <td>No other material provided</td>
                            <td>&mdash;</td>
                        </tr>
                    </tbody>
                </table> </table-wrap>
        </sec>
<fn-group>
    <fn fn-type="financial-disclosure"><p>This work was made possible through funding from the Institute of Education Sciences, U.S. Department of Education, Grant R305D210044.</p></fn>
</fn-group>
<fn-group>
    <fn fn-type="conflict"><p>The author has declared that no competing interests exist. The author wrote the paper himself and excluded AI.</p></fn>
</fn-group>
<ack>   
<p>The author has no additional (i.e., non-financial) support to report.</p>
</ack>
</back>
</article>
