
Changelog
TOSTER v0.9.0
New Features
New
trans_rank_prob()function for transforming probability-scale effect sizes between four scales: probability (concordance), difference (rank-biserial), log-odds, and odds. Supports bidirectional transformation viafromandtoarguments with delta-method standard errors and monotonic CI endpoint mapping.brunner_munzel()gains ascaleargument to report results on alternative scales (“probability”, “difference”, “logodds”, “odds”) without changing the underlying test. The defaultscale = "probability"preserves existing behavior.ses_calc()estimate labels now use probability notation (e.g.,P(X>Y) - P(X<Y)instead ofRank-Biserial Correlation). The$methodstring and data frame row names retain human-readable names. Code that parsesnames(result$estimate)may need updating.ses_calc()paired sample labels useP(Z>0)notation (where Z = X - Y); one-sample labels useP(X>0).-
Effect size calculators now support hypothesis testing and
htestoutput:-
ses_calcandboot_ses_calcupdated withoutput,alternative, andnull.valuearguments- Default output is now
"htest"class; useoutput = "data.frame"for legacy format - Supports
"two.sided","less","greater","equivalence", and"minimal.effect"alternatives - New Agresti/Lehmann placement-based SE method (
se_method = "agresti") with log-odds scale hypothesis testing - Continuity correction for boundary cases (complete separation)
- Default output is now
-
smd_calcandboot_smd_calcupdated withoutput,alternative,null.value, andtest_methodarguments- Default output is now
"htest"class; useoutput = "data.frame"for legacy format - Supports the same alternative hypothesis options as
ses_calc -
test_methodargument ("z"or"t") controls the reference distribution forsmd_calc - Degrees of freedom included in output when
test_method = "t" - Bootstrap p-values for
boot_smd_calccomputed from empirical distribution - New
denomargument for direct denominator selection ("z","rm","pooled","avg","glass1","glass2"), overridingglass,rm_correction, andvar.equalas needed; informative messages on conflicts
- Default output is now
-
-
Added
hodges_lehmannfunction for robust location testing- Implements Hodges-Lehmann estimators (HL1 for one-sample/paired, HL2 for two-sample)
- Supports exact permutation, randomization, and asymptotic (KDE-based) inference
- Full support for equivalence and minimal effect testing
- Consistent sign convention with
wilcox.test(x - y for two-sample tests) - Formula and default methods available
Added
perm_t_testfunction to allow for permutation tests for equivalence using TOST-
Added
plot_htest_est()function to create simple estimate plots from anyhtestobject- Displays point estimate with confidence interval
- Handles null values (single or equivalence bounds) as reference lines
- Automatically handles two-sample t-test estimates by computing mean difference
-
Added BCa (bias-corrected and accelerated) bootstrap confidence intervals as a new
boot_ci = "bca"option for:-
boot_t_test,boot_t_TOST,boot_log_TOST,boot_smd_calc,boot_ses_calc, andboot_cor_test - BCa intervals provide second-order accuracy by correcting for bias and skewness in the bootstrap distribution
- Acceleration factor computed via leave-one-out jackknife (pooled jackknife for two-sample designs)
- Informative errors for degenerate cases with suggestion to use
boot_ci = "perc"as fallback
-
Improvements
boot_smd_calc()now defaults toboot_ci = "bca"(previously"stud"). In simulations with skewed data, the studentized interval was liberal (two-sided Type I error of about 0.10 to 0.11 at a nominal 0.05 with 20 to 50 observations per group), because its pivot uses a normal-theory standard error for the SMD. BCa stayed at or below the nominal rate in all conditions studied. Results for code that relied on the default will change; setboot_ci = "stud"to reproduce earlier output.-
Correlation SE improvements for
z_cor_test()andcorsum_test():- Spearman’s rho now uses the Bonett-Wright ρ-dependent SE formula (
sqrt((1 + r^2/2) / (n - 3))) instead of the fixed 1.06 constant, providing better calibration across the full range of rho. -
z_cor_test()gains ase_methodargument ("analytic"or"jackknife") for computing the standard error via leave-one-out resampling on the Fisher z scale. The jackknife SE is used consistently for both the test statistic and the confidence interval. - Both functions now return
stderras a named vector withz.se(Fisher z scale, used for inference) andcor.se(delta method SE on the correlation scale, for descriptive purposes). - The
methodstring in the returnedhtestobject now indicates the SE type used (e.g.,"Pearson's product-moment correlation with approximate SE"or"Spearman's rank correlation rho with jackknifed SE").
- Spearman’s rho now uses the Bonett-Wright ρ-dependent SE formula (
simple_htest(),boot_t_test(),perm_t_test(), andhodges_lehmann()now produce more informative estimate labels that indicate the direction of calculation (e.g.,"mean difference (treatment - control)"when using the formula interface).For two-sample mean-based tests (
simple_htestwith t-test,boot_t_test,perm_t_test), the mean difference is now appended as a third element of$estimate, while preserving existing group means at positions 1-2 (backwards-compatible).Paired test estimates are labeled to clarify the differencing operation, e.g.,
"mean of the differences (z = x - y)".Wilcoxon/Mann-Whitney estimates in
simple_htest()are labeled as"Hodges-Lehmann estimate"with direction indicated.hodges_lehmann()estimate labels updated for clarity:"pseudomedian of x"for one-sample,"Hodges-Lehmann estimate (x - y)"for two-sample, and"Hodges-Lehmann estimate (z = x - y)"for paired tests.Trimmed mean labels in
boot_t_test()andperm_t_test()now include the trimming proportion, e.g.,"trimmed mean difference (x - y, tr = 0.1)".A
$sample_sizeelement (named numeric vector) is now included in the returnedhtestobject for all four functions. For two-sample formula calls, names reflect the actual factor levels.Permutation test terminology: Clarified distinction between “Exact Permutation” (all permutations enumerated) and “Randomization” (permutations sampled with replacement) tests across
perm_t_test,hodges_lehmann, andbrunner_munzel-
p_method auto-selection: Added intelligent default for
p_methodargument in permutation-based functions:-
NULL(default): Automatically selects “exact” for exact permutation tests and “plusone” for randomization tests - “exact”: Uses b/R, appropriate when all permutations are enumerated
- “plusone”: Uses (b+1)/(R+1) following Phipson & Smyth (2010), provides exact Type I error control for randomization tests
-
Update
brunner_munzelfunction to allow TOST directlyUpdate functions to disallow
paired = TRUEwhen formula method utilized.-
Improved
plot.TOSTtfortype = "simple":- Raw estimate plot now appears on top (was on bottom)
- Decision text and equivalence bounds now displayed at top of plot
- Added
layoutparameter: “stacked” (default) or “combined” for a single faceted plot
-
Improved
plot.TOSTtfortype = "tnull":- Now shows only one-sided rejection regions appropriate to the test type
- Equivalence tests: lower bound shows right tail, upper bound shows left tail
- Minimal effect tests: lower bound shows left tail, upper bound shows right tail
perm_t_test()andboot_t_test()documentation now describes the sharp (Fisher) versus weak (Neyman) null hypotheses: the permutation test is exact under the sharp null and, when studentized, asymptotically valid for the weak null (Wu & Ding, 2020), while the bootstrap targets the weak null and is asymptotic only. Guidance on randomized versus random-sampling designs and on small-sample, unequal-variance behavior is included.perm_t_test(),boot_t_test(),boot_t_TOST(), and the robust TOST vignette now include a “Choosing a Method” guide based on a set of simulations: studentized permutation for independent groups; the studentized bootstrap for paired data with clearly skewed differences (the sign-flip permutation test assumes symmetry); trimming for heavy tails or outliers; caution when groups differ in shape; andboot_ci = "bca"is not recommended for mean differences because it tended to be too liberal.brunner_munzel()gainstest_method = "perm_logit": a studentized permutation test on the logit scale. Its confidence interval inverts the same test and is back-transformed, so it is range-preserving (never clamped) and always agrees with the p-value. In simulations it gave the most powerful equivalence tests while keeping Type I error near nominal, and intervals that do not collapse when the estimate is near 0 or 1. It is the recommended method for equivalence, minimal effect, and other tests against a null value other than 0.5.brunner_munzel()now warns whentest_method = "t"or"perm"is used for a minimal effect test or any other test against a null value other than 0.5. In simulations these methods had inflated Type I error for minimal effect tests (up to about 15–18% in small samples with wide bounds) and low power for equivalence tests. The warning suggests"perm_logit"or"logit".-
brunner_munzel()test method guidance is updated based on simulation studies of Type I error (see the new “Choosing a test method” section):- Two-sample: the message recommending
test_method = "perm"now appears when the smaller group has fewer than 30 observations (previously 15). - Paired:
"perm"is now recommended for any sample size, because the"t"and"logit"methods are conservative when pairs are positively correlated. The “permutation test is probably unnecessary” message is now shown only for two-sample designs. - A new message warns when the permutation distribution is too coarse for the test to ever reject at the requested
alpha(e.g., 5 or fewer pairs).
- Two-sample: the message recommending
Bug Fixes
Two-sample bootstrap in
boot_smd_calc()andboot_ses_calc()now resamples within each group. Previously, observations were resampled from the pooled data, so group sizes varied across bootstrap replicates. With small samples a group could be left with one or zero observations, producingNaNestimates and anNAp-value or an error (about 4% of calls with 10 observations per group). Group sizes are now fixed at the observedn1andn2, matchingboot_t_test()andboot_t_TOST().-
Consistent handling of
muint_TOST(),tsum_TOST(), andboot_t_TOST():- The raw estimate, its confidence interval, and the raw equivalence bounds are now all reported on the original scale. Previously the raw estimate was
estimate - muwhile the confidence interval and bounds were not shifted. TOST p-values fromt_TOST()andtsum_TOST()are unchanged. - The SMD and its bounds are now consistently relative to
mu(e.g.,(x - y - mu) / SD). Previously the two-sample SMD addedmu, the paired SMD ignoredmu, andtsum_TOST()ignoredmufor all designs. Bounds given witheqbound_type = "SMD"are standardized distances frommu. -
muis now stored in the returnedTOSTtobject.print()reports the equivalence bounds and notes the scale of each row whenmuis not zero, anddescribe()uses the storedmu(previously always 0 fortsum_TOST()). - The “Equivalence interval does not include zero” message now checks whether the bounds contain
mu.
- The raw estimate, its confidence interval, and the raw equivalence bounds are now all reported on the original scale. Previously the raw estimate was
-
boot_t_TOST(): the studentized bootstrap p-values used a bootstrap t-statistic whose variance was centered incorrectly, which under-dispersed the reference distribution whenever the (difference in) means was far from zero. The p-values now use the same pivot as the studentized confidence interval, so they agree with the interval and withboot_t_test().- Paired resamples are now drawn in the same order as
boot_t_test(), so both functions give identical p-values and confidence intervals for the same seed (results for a given seed differ from earlier versions). - The Welch two-sample bootstrap replicates now use the normal-approximation SMD standard error, as the other designs already did.
- Paired resamples are now drawn in the same order as
boot_t_test()withvar.equal = TRUE: the studentized confidence interval used Welch standard errors for the bootstrap replicates while the observed standard error and p-value used the pooled standard error. The replicates now use the pooled standard error, so the interval and p-value agree.-
boot_cor_test(boot_ci = "stud")was not actually studentized. The pivots used normal-theory Fisher z standard errors that depend only onn(Pearson, Kendall) or on the estimate itself (Spearman), so the interval was in effect a basic bootstrap interval on the z scale. Each bootstrap replicate is now standardized by an influence-function (sandwich) standard error estimated from that replicate’s data, which does not assume bivariate normality:- Pearson: asymptotic distribution-free (fourth-moment) standard error with an HC4-type leverage correction, which keeps coverage near nominal for small samples and heavy-tailed data.
- Spearman: influence function of Pearson’s r on the mid-distribution transforms (midranks), including the terms for estimating the transforms, so it stays accurate with ties.
- Kendall: U-statistic (Hoeffding projection) standard error of tau-b. The formulas are given in the new “Studentized bootstrap” section of
?boot_cor_testand invignette("correlations"). Studentized results, and thez.seelement ofstderr, will differ from earlier versions.
boot_cor_test()has a new default,boot_ci = "auto", which uses the studentized interval ("stud") for Pearson’s r and BCa ("bca") for the Spearman, Kendall, Winsorized, and percentage bend correlations. In simulations, the studentized Pearson interval stayed near nominal coverage for skewed, heteroscedastic, and heavy-tailed data, where BCa under-covered. Results for Pearson correlations with the default settings will differ from earlier versions (useboot_ci = "bca"for the previous default). The returnedboot_cielement reports the method that was used.boot_cor_test()gains aboot_scaleargument ("z", the default, or"r") setting the scale on which the"basic"and"stud"intervals and p-values are computed. The basic interval was previously computed on the correlation scale while the studentized interval used the Fisher z scale; both now default to the z scale, so"basic"results will differ from earlier versions (useboot_scale = "r"for the previous basic interval). The"perc"and"bca"methods are unaffected. The result also gains aboot_scaleelement.perm_t_test(): the confidence interval was a percentile interval of the raw (non-studentized) permuted differences, while the p-value came from the studentized permutation test. Under unequal variances and group sizes the two could disagree, and the percentile interval was also reflected in the wrong direction for asymmetric permutation distributions (#120). The confidence interval is now obtained by inverting the same permutation test (samep_methodcounting rule andsymmetricsetting), so it always agrees with the p-value. Confidence interval values will differ from earlier versions.hodges_lehmann()permutation tests: the confidence interval was likewise a percentile interval of the raw permuted estimates (shifted by the estimate) rather than an inversion of the permutation test, so it could disagree with the p-value. It now inverts the same test (samep_methodcounting rule and absolute-value two-sided rule), matching theperm_t_test()fix (#120).brunner_munzel(test_method = "perm"): the confidence interval now inverts the same studentized permutation test as the p-value (#120). Previously the two-sample two-sided interval was equal-tailed while the p-value used the absolute-value rule, the equivalence/minimal effect interval mirrored the upper quantile to both sides, the paired intervals assumed a symmetric permutation distribution, and the order-statistic rules did not match thep_methodcounting rule. The permutation minimal effect p-value now counts the opposite tails directly instead of using1 - p.Permutation p-values in
perm_t_test(),brunner_munzel(test_method = "perm"), andhodges_lehmann()now treat permutation statistics within floating point error of the observed statistic as ties. With tied data (or, forbrunner_munzel(), whenever the observed and permuted statistics are computed by different code paths), mathematically tied permutations could differ in the last bits and be dropped from the count, making p-values too small. For example, the exact Brunner-Munzel permutation test with 5 vs 5 ordinal data rejected about 8% of the time under exchangeability (now about 1.4%), and fully enumerated tests could return p = 0, which is impossible when the observed arrangement is among the permutations. Confidence intervals use the same tolerance, so they continue to agree with the p-values.smd_calc()andboot_smd_calc(): the one-sample SMD now usesmean(x) - mucorrectlyjamovi one-sample TOST now passes the
muoption to the analysis.plot.TOSTt(type = "tnull"): fixed swapped internal labels for the CI limits.
TOSTER v0.8.7
- Update documentation to make it clear what the “eqb” argument does within the
wilcox_TOSTfunction.
TOSTER v0.8.6
CRAN release: 2025-08-22
- Add warning message about error control with MET
- Fix unit tests for boot_ses_calc to catch errors when estimates contain infinite values or when ses is not “rb”
TOSTER v0.8.5
- Big update to package documentation to make things more detailed.
- Added extra message when permutation tests are used for the Brunner-Munzel test.
- Added more alternative hypotheses to
boot_compare_smd. - Expanded functionality of
power_eq_ffunction.
TOSTER v0.8.4
CRAN release: 2025-02-06
- Added simple plot for
TOSTtmethods - Small fix to output for printed method for
TOSTt - Added fix to
power_z_corto provide “zero” power for scenarios where the effect is undetectable
TOSTER v0.8.3
CRAN release: 2024-05-08
- Change in the standard error formulation to Glass delta for independent samples
- Hat tip to Paul Dudgeon for catching an error in the code that led to this development
- Other small changes to standard error calcs (see vignettes for details)
- Small modification to
plot_smdto catch errors when attempting to plot - Brunner-Munzel updates
- Added more warning messages
- Changed default for
simple_htestto 0.5 rather than 0
TOSTER v0.8.2
CRAN release: 2024-04-16
- Fixed error with
describemethod for minimal effects test forTOSTtobjects.
TOSTER v0.8.1
CRAN release: 2024-03-21
- Small correction to the displayed equation for Cohen’s ds standard error. Thank you to Matthew B Jané for finding this error.
- Added bootstrap options such as
boot_smd_calcandboot_ses_calc.- Many functions also now allow for different CI methods for bootstrapped results.
TOSTER v0.8.0
CRAN release: 2023-09-14
- Added Brunner-Munzel test
- Updated documentation to include lifecycle labels
- Created new function for two proportions tests (
twoprop_test)- And power
power_twoprop
- And power
- Created new function for power for correlations (
power_z_cor) - Deprecated old functions
TOSTER v0.7.1
CRAN release: 2023-04-05
- Fixing the ggplot2 error related to
after_statupdate from that package. - Update documentation.
TOSTER v0.7.0
- Add correlation functions
z_cor_test,boot_cor_test,corsum_test,boot_compare_cor, andsimple_htest - Add
describemethod to provide verbose output for analyses- Alternative
describe_htestfunction for htest objects
- Alternative
TOSTER v0.6.0
CRAN release: 2022-12-13
- Changed Glass’s delta SE for paired samples (minor).
- Added
smd_calcandses_calcfor just calculating the standardized effect sizes (no tests). - Default CIs for for SMDs are now NCT rather than the Goulet method.
-
compare_smdcan be supplied with user provided standard errors. - Add
log_TOSTandboot_log_TOSTfunction for comparing ratios of means. - Reduce the amount of text in the
printmethods
TOSTER v0.5.0
- Added “compare” functions.
-
compare_smd: Compare 2 SMDs from summary statistics -
boot_compare_smd: Compare 2 SMDs from raw data -
compare_cor: Compare 2 independent correlations
-
- Added additional SMD options
- Confidence intervals can now be estimated using other methods
- smd_ci can be used to set the confidence interval method
- Glass’s delta can now be calculated using the
glassargument
- Added additional standardized effect sizes for
wilcox_TOST-
sesargument can be set to “r”, “odds”, or “cstat” - Respectively, these will provide the rank-biserial correlation, odds, or concordance probability
-
TOSTER v0.4.0
CRAN release: 2022-02-04
“Avocado TOST” (Release date: 2022-02-05)
Changes:
-
New t_TOST function
- Generalized function for TOST for any type of t-test
- S3 generic methods to print and plot results
-
New power_t_TOST function
- Generalized function for t_TOST power analysis
- Outputs the an object of
power.htestclass
- Updated all jamovi functions to allow minimal effect tests
- Direction of one-sided tests now allows
- Added equ_anova and equ_ftest
- Now allows equivalence (also called non-inferiority)
- jamovi functions using t-test have more plotting options
- Error in powerTOSTtwo fixed when determining N
- All old t-test based TOST functions now have message telling users they are defunct
- All TOST procedures have text results changed to provide more appropriate feedback.
TOSTER v0.3.4
CRAN release: 2018-08-03
(Release date: 2018-08-05)
Changes:
- Added a verbose = FALSE option to all functions.
- Added Fisher’s z transformed CI to output of TOSTr.
- Cleaned up the text and numeric output of the functions.
- Renamed some variable names in the TOSTtwo.prop function for similarity with TOSTmeta.
- Added warnings and error messages for possible incorrect input when using the functions.
TOSTER v0.3.3
CRAN release: 2018-05-08
(Release date: 2018-05-08)
Changes:
- Error in order in which p-value 1 and p-value 2 were reported for TOSTr due to incorrect use of abs(r) in function. (thanks to Nils Kroemer and Dan Quintana)
- Fixed error in powerTOSTr function - instead of based on conversion to d, new function calculates directly from r and is more accurate.
- Added testthat folder and initial unit tests.
TOSTER v0.3.2
CRAN release: 2018-04-14
(Release date: 2018-04-14)
Changes:
- Error in sensitivity analyses of powerTOSTpaired.raw, powerTOSTone.raw, powerTOSTtwo.raw where output was in Cohen’s d - now changed to output in raw scores, added sd to examples. (thanks to Lisa DeBruine)
- Raw power functions still output ceiling N in message, but exact N in output value
TOSTER v0.3.1
CRAN release: 2018-04-06
(Release date: 2018-04-06)
Changes:
- Error in TOSTpaired.raw where t-test for TOST multiplied by sdiff (copied from TOSTpaired function) removed. CI were correct, but p-values did not match. Now tests are correct. (thanks to ontogenerator)
- Changed text in TOSTr function (text copied from t-test script now changed to correlation)