
Multiple Imputation and Parallel Workflows
Source:vignettes/qgcompmulti-mi-parallel.Rmd
qgcompmulti-mi-parallel.RmdIntroduction
The package includes two workflow-oriented capabilities that often matter together in practice:
- native multiple-imputation fitting through
qgcomp.glm.multi.mi() - optional bootstrap-level parallel execution for ordinary and MI workflows
This article is a short companion to the main workflow vignette. It focuses on when to use the native MI wrapper, what the pooled object contains, and how the parallel option fits into larger repeated-fit analyses.
When to use the native MI wrapper
Use qgcomp.glm.multi.mi() when you want
qgcomp.multi to:
- fit one
qgcomp.glm.multi()model per imputation - pool the MSM coefficients and covariance internally using Rubin’s rules
The pooled object is focused on inference and summary. It does not
try to reproduce the full prediction and diagnostics surface of a
single-fit qgcompmulti object.
A small completed-data example
library(qgcomp.multi)
dat <- sim_mixture_data(
n = 120,
pA = 3,
pB = 3,
rho_within_A = 0.3,
rho_within_B = 0.3,
rho_between = 0.2,
psi1 = 0.5,
psi2 = 0.3,
psi12 = 0.2,
return_quantized = FALSE,
seed = 123
)
imp_list <- lapply(
seq_len(2),
function(i) {
dat_i <- dat
dat_i$X1 <- dat_i$X1 + rnorm(nrow(dat_i), sd = 0.02 * i)
dat_i$W2 <- dat_i$W2 + rnorm(nrow(dat_i), sd = 0.02 * i)
dat_i
}
)Fit the pooled model:
fit_mi <- qgcomp.glm.multi.mi(
f = Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C,
data = imp_list,
mix1 = c("X1", "X2", "X3"),
mix2 = c("W1", "W2", "W3"),
interaction = TRUE,
q = 4,
B = 5,
seed = 13
)Inspect the pooled result:
summary(fit_mi)
#> Summary of qgcompmulti multiple-imputation fit
#>
#> Call:
#> qgcomp.glm.multi.mi(f = Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C,
#> data = imp_list, mix1 = c("X1", "X2", "X3"), mix2 = c("W1",
#> "W2", "W3"), interaction = TRUE, q = 4, B = 5, seed = 13)
#>
#> Model overview:
#> Formula: Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C
#> Outcome: Y
#> Family: gaussian (identity)
#> Estimand: Mean difference (default)
#> MSM fitting scale: identity
#> Interval method: wald (pooled multiple imputation)
#> Observations used: 120
#> Exposure mode: Quantized exposures (q = 4)
#> MSM interaction: included
#> Bootstrap replications per imputation: 5
#> Monte Carlo size: 120
#>
#> Multiple imputation overview:
#> Imputed datasets: 2
#> Input type: Completed data list
#> Retained per-imputation fits: no
#> Master seed: 13
#> Stored fit-specific seeds: 2
#>
#> Mixtures:
#> Mixture 1: X1, X2, X3
#> Mixture 2: W1, W2, W3
#>
#> MSM coefficients:
#> Estimate Std. Error t value df Pr(>|t|)
#> Intercept -0.36260 0.58396 -0.62094 2.6315e+04 0.534645
#> Mixture 1 main effect 1.00692 0.36985 2.72250 1.2986e+03 0.006566 **
#> Mixture 2 main effect 0.69541 0.38841 1.79043 3.9763e+10 0.073385 .
#> Mixture interaction -0.13804 0.18639 -0.74063 7.6298e+03 0.458939
#> ---
#> Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
#>
#> Pooling diagnostics:
#> df RIV FMI
#> Intercept 2.6315e+04 6.2028e-03 0.0062
#> Mixture 1 main effect 1.2986e+03 2.8542e-02 0.0278
#> Mixture 2 main effect 3.9763e+10 5.0149e-06 0.0000
#> Mixture interaction 7.6298e+03 1.1581e-02 0.0114
coef(fit_mi)
#> (Intercept) psi1 psi2 psi1:psi2
#> -0.3626040 1.0069164 0.6954136 -0.1380450
confint(fit_mi)
#> 2.5 % 97.5 %
#> (Intercept) -1.50719699 0.7819890
#> psi1 0.28134774 1.7324851
#> psi2 -0.06584764 1.4566748
#> psi1:psi2 -0.50341647 0.2273265Retaining per-imputation fits
If you want the full completed-data fits in addition to the pooled
result, set keep_fits = TRUE.
fit_mi_keep <- qgcomp.glm.multi.mi(
f = Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C,
data = imp_list,
mix1 = c("X1", "X2", "X3"),
mix2 = c("W1", "W2", "W3"),
interaction = TRUE,
q = NULL,
centering = "median",
B = 5,
seed = 17,
keep_fits = TRUE
)
fit_mi_keep$mi_info$keep_fits
fit_mi_keep$fits$imputation_fits[[1]]This is useful when you want pooled inference as the main result but
still need to inspect one or more single-fit qgcompmulti
objects directly.
mids objects are also supported
The wrapper also accepts mice::mids objects when the
optional mice package is installed.
library(mice)
dat_miss <- dat
dat_miss$X1[sample.int(nrow(dat_miss), 8)] <- NA_real_
dat_miss$W2[sample.int(nrow(dat_miss), 8)] <- NA_real_
mids_obj <- mice(dat_miss, m = 3, maxit = 1, printFlag = FALSE)
fit_mi_mids <- qgcomp.glm.multi.mi(
f = Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C,
data = mids_obj,
mix1 = c("X1", "X2", "X3"),
mix2 = c("W1", "W2", "W3"),
interaction = TRUE,
q = 4,
B = 5,
seed = 19
)Optional bootstrap-level parallelism
The package supports one level of optional parallelism.
For ordinary fits, bootstrap replications can be dispatched in
parallel inside qgcomp.glm.multi():
fit_parallel <- qgcomp.glm.multi(
f = Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C,
data = dat,
mix1 = c("X1", "X2", "X3"),
mix2 = c("W1", "W2", "W3"),
interaction = TRUE,
q = 4,
B = 20,
seed = 29,
parallel = TRUE,
workers = 2
)For native MI fits, the imputation loop stays serial, but the bootstrap work inside each completed-data fit can still be parallelized:
fit_mi_parallel <- qgcomp.glm.multi.mi(
f = Y ~ X1 + X2 + X3 + W1 + W2 + W3 + C,
data = imp_list,
mix1 = c("X1", "X2", "X3"),
mix2 = c("W1", "W2", "W3"),
interaction = TRUE,
q = 4,
B = 10,
seed = 31,
parallel = TRUE,
workers = 2
)Practical notes
Several practical limits are worth stating plainly:
- the MI loop itself is still serial
- reproducibility is defined within a fixed backend and execution mode, not as exact equality between serial and parallel runs
-
progress = TRUEis supported only in serial mode and is disabled with a clear warning whenparallel = TRUE - pooled prediction, pooled diagnostics, pooled bootstrap interval methods, and pooled confidence regions are not yet the main target of the pooled MI object
In practice, the native MI wrapper is best used when pooled MSM inference is the main goal and you want the package to manage the repeated fitting and Rubin pooling directly.