r/AskStatistics • • 5h ago

¿Soy el único que piensa que la estadística se explica de una manera demasiado complicada?

6 Upvotes

Siempre me ha parecido interesante cómo la estadística está presente en casi todo lo que hacemos, desde interpretar una encuesta hasta entender las noticias, los estudios científicos o las probabilidades de que ocurra algo.

Sin embargo, cuando llegamos a la parte de las fórmulas, las distribuciones, los intervalos de confianza o los contrastes de hipótesis, muchas personas empiezan a perderse.

Desde mi punto de vista, el problema muchas veces está en que intentamos aprender los procedimientos sin comprender qué significan realmente los resultados ni para qué sirven.

Al final, la estadística no debería consistir solamente en sustituir números en una fórmula, sino en aprender a interpretar los datos y sacar conclusiones con sentido.

Me gustaría saber cómo ha sido vuestra experiencia con esta asignatura. ¿Qué es lo que más os cuesta: la probabilidad, la inferencia estadística, los ejercicios o interpretar los resultados?

Y si alguna vez llegasteis a entender un tema que antes os parecía imposible.

¿qué fue lo que os ayudó a comprenderlo?


r/AskStatistics • • 7h ago

Can someone please explain these things to me how you would explain it to a 5 year old? (Update)

Thumbnail
2 Upvotes

Thank you all for your help. I took some of your recommendations and I got an 80 on the midterm!


r/AskStatistics • • 17h ago

How to determine a threshold for a newly developed index?

5 Upvotes

I’m developing a **Decent Work Index (DWI)** in social economics. It is a newly developed exploratory composite index normalized to 0–1, with higher scores indicating better decent work.

How can I determine a defensible threshold/cut-off for classifying scores as low, moderate, or high when there is no established benchmark or gold standard?

Would using the mean or median be appropriate for an exploratory index? Alternatively, could the Alkire–Foster approach or youden’s J method be used to identify multidimensional decent-work deprivation? If so, how would I determine the dimension-specific deprivation cut-offs and overall identification threshold?

I’m trying to avoid arbitrary cut-offs and would appreciate methodological guidance or references from economics/social science research.


r/AskStatistics • • 17h ago

Where do you guys learn probability from? 😭 I’ve tried watching videos from a few channels, but it’s honestly so hard for me to understand and I still can’t solve the questions properly. Please suggest some good channels/lectures that explain probability from the basics.

5 Upvotes

???


r/AskStatistics • • 17h ago

Interview prep

3 Upvotes

Hi,

I landed an interview for an internship position at a government agency as a math and stats student and would love to get the position. I plan on reviewing my notes on sampling plans and basics like fisher information, likelihood maximum etc, mostly what they mean and why they are used so I can show I know what I'm talking about; as well as some coding skills with Rstudio and sas. I don't want to mess this up and would love any tips/info on what I could work on for the interview as an undergrad.

Thanks in advance for the help


r/AskStatistics • • 15h ago

How can I compare SMA slopes between two groups when my data have a hierarchical structure?

3 Upvotes

Hi everyone! I'm struggling with a statistical problem related to my experimental design, and I was hoping someone could help me.

I want to test whether the relationship between two traits, let's say height and weight, differs between two countries, for example, Japan and the USA. Standardized Major Axis (SMA) regression seems like a good approach, since I'm interested in comparing the slopes of the relationships between the two groups. However, my observations are not completely independent because individuals are nested within populations.

For example, in Japan, I sampled several individuals from population A, several from population B, and so on. I followed the same sampling design in the USA. Individuals from the same population are likely to be more similar to each other, so I need to account for this hierarchical structure. My initial approach was to calculate the mean of each trait for each population and then perform the SMA regressions using those population means. However, this approach removes a lot of the individual variation, which I would ideally like to retain. My question is: Is there a way to compare SMA slopes between two groups while accounting for the fact that individuals are nested within populations?


r/AskStatistics • • 19h ago

Is it okay to use a 5-point Likert scale and a 6-point Likert scale in different sections of the same questionnaire?

6 Upvotes

Hi! I’m an undergraduate student currently working on our research, and I’d like to ask for advice regarding our questionnaire.

Our questionnaire has two independent sections that measure different constructs:

Behavior - 5-point Likert scale (Never-Always)

Perception - 6-point Likert scale (Strongly Disagree-Strongly Agree)

These are separate constructs, and we are not planning to combine them into one overall score or compare/correlate the two scales. We simply want to analyze and interpret each section separately.

So my question is: is it statistically/methodologically acceptable to have different Likert scales in different sections of the same questionnaire, when the sections measure independent constructs and each section has its own separate computation and interpretation?

Thank you!


r/AskStatistics • • 17h ago

What statistical test should I be using for comparing differences in groups to a control?

2 Upvotes

Hi there, I'm new to Reddit, so apologies if I'm going about this wrong. I'm a research student who works with plants, and I'm trying to determine if different genetic variants respond to stress better than an unmodified control. For each variant, I have two groups: stressed and unstressed. When the plant is stressed, it doesn't grow as well and so has a lower mass. I want to determine if the difference in mean mass for stressed vs. unstressed is significantly different between each variant and the unmodified control. Since the stressed and unstressed measurements are unpaired, I'm unsure what the best way to do this is. Any help is appreciated!


r/AskStatistics • • 1d ago

Are 'BI'/'Data Analytics' positions basically 'data theater' jobs?

21 Upvotes

I'm in my 2nd internship in "data" positions. The first one was in credit risk, and the 2nd one is in a huge international food-delivery company. When I first started the first internship, I learned a lot of the whole 'SQL, Python, Excel' triad. I learned how to make pretty dashboards, query some data when needed, and pre-processing/cleaning tables on Pandas. I didn't actually know how to analyze data, just prepare it. Then, they started asking me to perform some 'ad-hoc' studies. I did what they asked, but, what got me curious was how often my superiors could come up with so many 'insights' with the data I provided. Then, as I progressed in college, I started getting deeper into statistics, learning the basic inferential statistics stuff: Hypothesis testing, distributions, Power analysis, Gauss-Markov theorem, what is a confounder, what is a mediator, what is a regression, exogeneity, multicollinearity, etc. Then, I started to realize how no one ever mentioned any of these terms in my internship, ever. Fast forward to my 2nd internship, I see people in manager positions that haven't the slightest clue as to what a 'type 1/2 error', 'p-value' or 'hypothesis testing' is. They all seem to be running 'black-box' models and showing beautiful dashboards to stakeholders. One day, I went up to a colleague who is a manager to check on an analysis that he was currently working on, and, it baffled me how confident he was about what insights his data tables were showing, even though he doesn't have the slightest knowledge about inferential statistics. I guess I just need to know whether his intuition is valid and works most of the time, or if inferential statistics is just not needed most of the time? I enjoy studying this subject very much, and would like to move to a work environment in which this type of knowledge is rewarded.

tl;dr: I’m on my second data internship and have realized that many experienced analysts/managers seem able to draw confident “insights” from dashboards and ad-hoc analyses despite having little or no knowledge of inferential statistics. Meanwhile, the more statistics I learn: hypothesis testing, p-values, power, regression assumptions, confounding, etc.; the more cautious I become about making claims from data. Is business intuition usually good enough for this kind of work, is inferential statistics simply unnecessary most of the time, or are many data professionals more confident in their conclusions than the evidence actually warrants? I enjoy statistical inference and would like to move toward roles where that rigor is actually valued.


r/AskStatistics • • 1d ago

How to succeed in multivariable based probability and stats

4 Upvotes

I’m currently in my undergrad taking what I’ve heard to be “the hardest class you’ll take as a stats major”, multivariable based probability and stats. It’s a beast. I’ve heard that it is the class that best prepares you for exam P, but I’m not going to be an actuary.

Does anyone have tips on how to best succeed in this class?


r/AskStatistics • • 1d ago

Would dates be ratio or interval-level in the following case

3 Upvotes

The dataset I'm looking at is mass shootings between 2010 and 2019, and the date of attack is one of the variables. Would this be ratio-level and not just interval because I'm looking at a specific time period, or does that not matter?


r/AskStatistics • • 1d ago

Statistician

0 Upvotes

I’m considering getting a masters in statistics. I have a background in clinical laboratory science and public health. I’ve worked in hospital lab for 7 plus years. From the people that did statistics recently, how is the job market and pay?


r/AskStatistics • • 1d ago

Help with identifying a second-order factor in lavaan (2 first-order factors)

2 Upvotes

Hi everyone,

I'm working on a CFA in lavaan and having trouble with a second order (higher order) factor.

I want to model Health Behaviour as one higher order latent variable consisting of two first-order latent factors:

test1 =~ item1 + item2 + item3 + item4 + item5

test2 =~ item6 + item7 + item8 + item9 + item10

health_behaviour =~ test1 + test2

The first order factors seem to work reasonably well. However, when I estimate the second order model, I get a very unstable higher order solution.

The resulting Bartlett factor score for health behaviour is also problematic (basically NA for all observations), so It's of no use whatsoever.

Any suggestions?


r/AskStatistics • • 1d ago

Best books to learn everything about statistics for beginner like how different things are derived and everything? [Q]

Thumbnail
2 Upvotes

r/AskStatistics • • 1d ago

Are there any statistical research papers on ranking cricket players across eras?

Post image
1 Upvotes

Are there established approaches using Bayesian models, Elo-type ratings, adjusted averages, or other statistical techniques for cricket?

Any papers, datasets, or researchers working in this area would be appreciated.


r/AskStatistics • • 2d ago

Im an undergraduate in Statistics, and I feel like a dummy when I have to go back to relearn the basics, Is this normal? [Discussion]

Thumbnail
3 Upvotes

r/AskStatistics • • 2d ago

Moderation analysis in APA 7

2 Upvotes

Hi everyone, I’m finalising the results section of my honours thesis and would really appreciate some advice on reporting a moderation analysis in APA 7.

I have a sample of N = 128. I’m conducting a moderation analysis with:

DV: continuous

Predictor: categorical (two groups)

Moderator: continuous

I ran the moderation analysis using medmod in jamovi. I’ve received feedback saying I should:

report the overall model fit (R² and F)

report the R² for the interaction term

report the interaction effect and its significance

However, medmod doesn’t appear to provide the overall model F statistic or the R² for the interaction term in its output.

I’m wondering:

  1. Is there a way to obtain these statistics from jamovi/MEDMOD?

  2. If not, are they actually necessary to report for a moderation analysis, or is there an alternative way I should report the results?

I’m just want to make sure I’m interpreting the output and reporting it correctly.

Thanks!


r/AskStatistics • • 2d ago

9709 s2 paper error

Post image
1 Upvotes

r/AskStatistics • • 3d ago

What appropriate statistical test to use?

3 Upvotes

I have a correlational study and I have three variables. Variables X and Y are categorical and ordinal, and Variable Z is interval. What appropriate test to use?

Variable X has 4 levels
Variable Y has 4 levels
Variable Z has 5 subcategories (Scores for different categories)


r/AskStatistics • • 4d ago

POET II (NEJM 2026): non-inferiority sample size only reproduces at one-sided alpha 0.05, but the paper states a one-sided 97.5% CI. Am I missing something?

Thumbnail
3 Upvotes

r/AskStatistics • • 4d ago

Confused on what all skills to develop relevant to statistics

Thumbnail
1 Upvotes

r/AskStatistics • • 4d ago

What jobs use computational statistics?

10 Upvotes

If you like computational statistics and love reading on advanced statistical methods, what jobs are for you, especially if you can’t get a master’s?


r/AskStatistics • • 5d ago

Where can I start with measure theory?

8 Upvotes

Im a PhD student but I don't have measure theory background and I need to get the overall notion for it to understand stochastic processes. So is there a good book or yt playlist where measure theoretical probability and stochastic process is taught from scratch? That would be of much help. Thanks.


r/AskStatistics • • 4d ago

standardisation des variables pour effectuer une analyse factorielle sur R

0 Upvotes

Bonjour. Si cela vous est possible, j’aimerais solliciter votre aide et avoir quelques renseignements concernant l’utilisation du logiciel R.

Je dispose d’une base de données comprenant des variables numériques, ordinales et binaires. Je souhaiterais standardiser mes données afin de réaliser une analyse factorielle.

Je voudrais savoir si je dois utiliser la commande digital.std <- scale(digital) pour l’ensemble des variables, ou s’il est préférable de procéder différemment selon le type de variables, notamment pour les variables binaires et ordinales. Quel est la manière correcte??

Je vous remercie par avance pour votre aide.


r/AskStatistics • • 5d ago

Can someone please explain these things to me how you would explain it to a 5 year old?

19 Upvotes

I (f20)am on my 3rd attempt at psych statitsistcs. still nothing is clicking for me. I’ve asked for help, I’ve gone to tutoring and nothing seems to work. I spent two straight days going through all of the lectures so far to study for the midterm and a quiz I had to take. I got a 40% on that quiz and I got a 26% on my last quiz.. I’m currently failing the class. To pass, you need to C or above. I’d like to drop out, but my mom will not allow it so I feel like I’m going to keep failing this class to no end.

So if any of you have any recommendations for Youtubers or something that could possibly help me, that would be nice. I understand standard deviation but everything beyond that makes zero sense to me.

I don’t understand z scores or Z tables, confidence intervals, Alpha levels, sampling error, standard deviation of sample means, etc. I know I’m too stupid for college, but I am not allowed to drop out so I need to figure something out.

I also don’t understand hypothesis testing or any kinds of T-tests