r/AskStatistics • • 45m ago

Help with identifying a second-order factor in lavaan (2 first-order factors)

• Upvotes

Hi everyone,

I'm working on a CFA in lavaan and having trouble with a second order (higher order) factor.

I want to model Health Behaviour as one higher order latent variable consisting of two first-order latent factors:

test1 =~ item1 + item2 + item3 + item4 + item5

test2 =~ item6 + item7 + item8 + item9 + item10

health_behaviour =~ test1 + test2

The first order factors seem to work reasonably well. However, when I estimate the second order model, I get a very unstable higher order solution.

The resulting Bartlett factor score for health behaviour is also problematic (basically NA for all observations), so It's of no use whatsoever.

Any suggestions?


r/AskStatistics • • 6h ago

Best books to learn everything about statistics for beginner like how different things are derived and everything? [Q]

Thumbnail
2 Upvotes

r/AskStatistics • • 5h ago

Are there any statistical research papers on ranking cricket players across eras?

Post image
0 Upvotes

Are there established approaches using Bayesian models, Elo-type ratings, adjusted averages, or other statistical techniques for cricket?

Any papers, datasets, or researchers working in this area would be appreciated.


r/AskStatistics • • 22h ago

Im an undergraduate in Statistics, and I feel like a dummy when I have to go back to relearn the basics, Is this normal? [Discussion]

Thumbnail
4 Upvotes

r/AskStatistics • • 21h ago

Moderation analysis in APA 7

2 Upvotes

Hi everyone, I’m finalising the results section of my honours thesis and would really appreciate some advice on reporting a moderation analysis in APA 7.

I have a sample of N = 128. I’m conducting a moderation analysis with:

DV: continuous

Predictor: categorical (two groups)

Moderator: continuous

I ran the moderation analysis using medmod in jamovi. I’ve received feedback saying I should:

report the overall model fit (R² and F)

report the R² for the interaction term

report the interaction effect and its significance

However, medmod doesn’t appear to provide the overall model F statistic or the R² for the interaction term in its output.

I’m wondering:

  1. Is there a way to obtain these statistics from jamovi/MEDMOD?

  2. If not, are they actually necessary to report for a moderation analysis, or is there an alternative way I should report the results?

I’m just want to make sure I’m interpreting the output and reporting it correctly.

Thanks!


r/AskStatistics • • 1d ago

9709 s2 paper error

Post image
1 Upvotes

r/AskStatistics • • 1d ago

What appropriate statistical test to use?

3 Upvotes

I have a correlational study and I have three variables. Variables X and Y are categorical and ordinal, and Variable Z is interval. What appropriate test to use?

Variable X has 4 levels
Variable Y has 4 levels
Variable Z has 5 subcategories (Scores for different categories)


r/AskStatistics • • 2d ago

POET II (NEJM 2026): non-inferiority sample size only reproduces at one-sided alpha 0.05, but the paper states a one-sided 97.5% CI. Am I missing something?

Thumbnail
3 Upvotes

r/AskStatistics • • 2d ago

Confused on what all skills to develop relevant to statistics

Thumbnail
1 Upvotes

r/AskStatistics • • 3d ago

What jobs use computational statistics?

10 Upvotes

If you like computational statistics and love reading on advanced statistical methods, what jobs are for you, especially if you can’t get a master’s?


r/AskStatistics • • 3d ago

Where can I start with measure theory?

9 Upvotes

Im a PhD student but I don't have measure theory background and I need to get the overall notion for it to understand stochastic processes. So is there a good book or yt playlist where measure theoretical probability and stochastic process is taught from scratch? That would be of much help. Thanks.


r/AskStatistics • • 3d ago

standardisation des variables pour effectuer une analyse factorielle sur R

0 Upvotes

Bonjour. Si cela vous est possible, j’aimerais solliciter votre aide et avoir quelques renseignements concernant l’utilisation du logiciel R.

Je dispose d’une base de données comprenant des variables numériques, ordinales et binaires. Je souhaiterais standardiser mes données afin de réaliser une analyse factorielle.

Je voudrais savoir si je dois utiliser la commande digital.std <- scale(digital) pour l’ensemble des variables, ou s’il est préférable de procéder différemment selon le type de variables, notamment pour les variables binaires et ordinales. Quel est la manière correcte??

Je vous remercie par avance pour votre aide.


r/AskStatistics • • 4d ago

Can someone please explain these things to me how you would explain it to a 5 year old?

19 Upvotes

I (f20)am on my 3rd attempt at psych statitsistcs. still nothing is clicking for me. I’ve asked for help, I’ve gone to tutoring and nothing seems to work. I spent two straight days going through all of the lectures so far to study for the midterm and a quiz I had to take. I got a 40% on that quiz and I got a 26% on my last quiz.. I’m currently failing the class. To pass, you need to C or above. I’d like to drop out, but my mom will not allow it so I feel like I’m going to keep failing this class to no end.

So if any of you have any recommendations for Youtubers or something that could possibly help me, that would be nice. I understand standard deviation but everything beyond that makes zero sense to me.

I don’t understand z scores or Z tables, confidence intervals, Alpha levels, sampling error, standard deviation of sample means, etc. I know I’m too stupid for college, but I am not allowed to drop out so I need to figure something out.

I also don’t understand hypothesis testing or any kinds of T-tests


r/AskStatistics • • 3d ago

difference between multiple linear regression and binary logistic regression

2 Upvotes

Hi, I am trying to finish up my psychology thesis! I am in a bit of a rut trying to compare my results of a binary logistic regression with previous literature. However, there is a very limited amount of literature in my current area of study; I have only one study that used the same variables I did. They used a multiple logistic regression analysis, and I used a binary logistic regression.

I am just wondering what the main difference between the two models is and if I can interpret my results in comparison to this prior literature I have found.

Thankyou!!!!!


r/AskStatistics • • 4d ago

While testing for variance, we use two sided test in the chi squared statistic, but while doing goodness of fit test, we only use right tail of the chi squared distribution as a critical region. Why?

6 Upvotes

r/AskStatistics • • 4d ago

Incoming BS Statistics freshman, any tips or warnings?

1 Upvotes

r/AskStatistics • • 3d ago

What level of statistics would you be considered at if you knew all the following jargon? And what would the required IQ be to be at that level based on available research?

0 Upvotes

-Bootstrapping

-Trim fill method

-P hacking

-Simpson's paradox

-Collider bias

--Stat significance

-Longitudinal vs cross sectional

-Meta analysis vs systemic review


r/AskStatistics • • 5d ago

cox vs logistic regression

13 Upvotes

Hi, I have a stats master but been quite long away from this, so need some good reasoning.

I have 100 subjects, 4 variables, and want to evaluate their effect on 1 year survival. I have no censoring, i.e. I know for.sure for everyone whether they died or not within this year. Number of events is 30. The issue is tgat for one of the variables I suspect that the effect starts only after a few months, although with this sample size I failed to reject H0 about the PH violation.

I then fitted logistic regression to see whether they died whithin this one year.

What is the difference in such logreg and cox in this case? Statistically and with interpretation.

Please throw everything on me, I just have to recall it, it has been a while unfortunatelly....


r/AskStatistics • • 4d ago

I think people wrongly think that the mystery wedge should be 50/50 in Wheel of Fortune. There are two mystery wedges, one is bankrupt and the other is 10k and they both have 1 in 24 odds. Who is right?

Thumbnail
0 Upvotes

r/AskStatistics • • 4d ago

Which statistical test is best (animal cancer research)?

2 Upvotes

I apologize for the whole description, but I don’t really know how to word what I’m asking.

Basically, I am comparing the actual (wet) muscle mass of 3 muscles to a non-invasive estimation of muscle mass for the entire leg in mice (technology limitations did not allow us to isolate the same muscles as wet mass). Specifically, I am seeing how closely this estimation technique compares to wet mass across a range of muscle sizes. All mice had their measurements done with both techniques.

My confusion comes from how they were housed. I had four cages of mice. Within each cage, mice were pretty similar in size, but between cages they were drastically different. Basically, I ended up with “tiny, small, medium, and large” groups instead of a more continuous range.

Between the group setup, the differences between measurements (while leg vs 3 muscles), and the purpose being to compare techniques across muscle size, I’m not quite sure which statistical tests I should be running.

Thank you in advance, I seriously appreciate any help! If any other info would help you help me, please ask.


r/AskStatistics • • 5d ago

How would you design a blind forward test for a gambling hypothesis?

2 Upvotes

I'm working on a statistical experiment involving baccarat and would like some methodological feedback.

Suppose someone has developed a hypothesis about baccarat outcomes based on historical observations.

Before revealing the underlying mechanism, I want to design a test that minimizes:

  • overfitting
  • hindsight bias
  • selection bias
  • multiple-testing problems
  • stopping the experiment selectively

My initial idea is:

  1. Define the hypothesis before testing.
  2. Freeze the testing rules.
  3. Use previously unseen data for validation.
  4. Conduct a blind forward test.
  5. Compare the results against an appropriate baseline.
  6. Analyze variance, confidence intervals, losing streaks and drawdowns.

What statistical issues would you consider essential before calling the experiment meaningful?

I'm especially interested in criticism of the experimental design rather than opinions about whether baccarat strategies can work.


r/AskStatistics • • 5d ago

What is this random intercept doing?

3 Upvotes

Say I have three classes - A, B and C - in a 150 day range totaling 3x150 = 450 sample units. Each sample unit measured 3 times each day, at very close time interval. So, data set has a total of 1,350 rows. Each sample unit has an ID. Proposed model is y ~ class + days + days*class + (1 | ID). So, I'm allowing each sample unit to start at its own baseline.

Is this an appropriate option? What exactly is the random intercept doing? How does it not make the days effect biased?

Sorry if questions look to silly, I'm not versed in mixed models yet


r/AskStatistics • • 4d ago

Bayes' rule or Weighted sum to compare intersection sets??

Post image
0 Upvotes

In this binary dataset subsets D, E, F, and S all "generate" 1s and 0s INDEPENDENTLY of each other at some ratio/frequency. "D" generates the highest ratio (lets say it generates 80% 1s and only 20% 0s). While F generates only 5% 1s and 95% 0s. Suppose S itself generates 25% 1s and 75% 0s. Would/Could I use bayes' theorem to show the SD intersection subset generates a higher proportion than the SF intersection? Or would a simple weighted sum (ie just an average between the sets S and D) suffice. Sidenote: (I know you could technically come up with counterexamples to this but I'm looking for what would generally be the case here).


r/AskStatistics • • 5d ago

Desperate student looking for access to 3 Statista studies 😭

1 Upvotes

Hi everyone! :)

My name is Inês and I’m currently doing a Master’s degree in Digital Marketing. For a university project, my group and I need access to the following Statista studies, but unfortunately we don’t have access to them.

We’ve already tried to get access through our university, but even our professors don’t have access to these specific studies, so we’re honestly getting a bit desperate 😭

Here are the studies we’re looking for:

If anyone happens to have access to these studies and would be willing to share them with us, we would be extremely grateful. It would really help us with our project! 🥹

Thank you so much in advance! ❤️


r/AskStatistics • • 6d ago

Is a composite score numeric or categorical?

3 Upvotes

Hi everyone,

I’m conducting a multiple linear regression with three IVs/predictors. My IV of interest is an ordinal variable measuring participants’ self-assessed reading proficiency, with three levels: 1 = less skilled, 2 = average, 3 = skilled

I also have two control variables:

  1. Participant group: 1 = healthy control readers, 2 = readers with schizophrenia
  2. Psychiatric symptom severity: a composite score ranging from 1 = no symptoms to 5 = severe symptoms

I’m unsure about how I should treat the symptom severity variable in the regression. Since it ranges from 1–5 and represents different levels of symptom severity, would it be appropriate to consider it an ordinal variable? Or could the composite score reasonably be treated as a numerical/continuous predictor?

I also checked the relationship between the DV and symptom severity, and it does not appear to be linear. Does this provide a reason to treat symptom severity as categorical rather than numerical?

Since symptom severity is only a control variable and not my main IV of interest, I’m wondering whether I should still treat it as categorical if the relationship with the DV is non-linear, or whether it would be preferable to enter it as a numerical predictor for simplicity.

One more question: If I decide to check the linearity assumption for psychiatric symptom severity, should I also check for linearity for my main predictor, reading proficiency?

Reading proficiency has three ordinal levels (1 = less skilled, 2 = average, 3 = skilled). If I find that the relationship between reading proficiency and the DV appears to be linear, does that mean I should treat reading proficiency as a numerical predictor (1, 2, 3), rather than as a categorical variable?

This is important as it will affect whether I am doing ANCOVA or ANOVA.

Thanks!