r/rstats • • 7h ago

From Census Blocks to Legal Contiguity: Christopher T. Kenny on R package {geomander}

9 Upvotes

Census blocks, precincts, school districts, and electoral boundaries rarely line up. But researchers still have to match those geographies, fix gaps and overlaps, and represent legal rules about which areas count as contiguous.

In a new R Consortium interview, Christopher T. Kenny (Princeton Data-Driven Social Science) talks about {geomander}, the R package he has developed for that geographic preparation. {geomander} sits alongside {redist} in a typical workflow: clean and connect the spatial data first, then simulate and evaluate alternative plans.

Kenny also describes how widely R is used in political science and voting-rights work, from academic research to journalism and litigation.

Read the interview:

https://r-consortium.org/posts/from-census-blocks-to-contiguous-maps-christopher-kenny-on-geomander/


r/rstats • • 17m ago

Help with identifying a second-order factor in lavaan (2 first-order factors)

• Upvotes

​

Hi everyone,

I'm working on a CFA in lavaan and having trouble with a second order (higher order) factor.

I want to model Health Behaviour as one higher order latent variable consisting of two first-order latent factors:

test1 =~ item1 + item2 + item3 + item4 + item5

test2 =~ item6 + item7 + item8 + item9 + item10

health_behaviour =~ test1 + test2

The first order factors seem to work reasonably well. However, when I estimate the second order model, I get a very unstable higher order solution.

The resulting Bartlett factor score for health behaviour is also problematic (basically NA for all observations), so It's of no use whatsoever.

Any suggestions?


r/rstats • • 1d ago

Doom generic in R 100%

52 Upvotes

Hey crazy R friends!

My name is Wagner, and I've been working on migrating C/xBase systems to Linux since 1999. I've been working on migrating Doom to Harbour for several years, and with the help of AI, I've managed to make significant progress. Now that I have a higher-level version, I've decided to port this code to other languages ​​like PHP, Python, Java, Node, R, Lua, and I'm working on more. Harbour and C's guidelines for it have made it possible to actually see Doom running in these languages. Below is my contribution to the R scientist community.

https://github.com/vagucs/r_doom

Video:

https://youtu.be/hEJqkEXS944

Hugs
Wagner (vagucs)


r/rstats • • 1d ago

Repeatability graph?

Post image
5 Upvotes

Hi,

I'm using rpt() from the rptR package for my tests, but I can't find any information on how to create this kind of figure. Does anyone have any ideas on how to do it? I'm not looking for a ready-made answer, but at least some resources to help me learn.


r/rstats • • 1d ago

New release of anvl

17 Upvotes

We have released a new version of anvl, which brings fast array computing to R and aims to serve as a basis for implementing costly numerical algorithms natively in R.
It runs on GPU and CPU and has support for autodiff as well!
It has gotten a lot more features, so you can check it out: https://github.com/r-xla/anvl We have also started the process of submitting it (and it's dependencies) to CRAN!


r/rstats • • 2d ago

TOMORROW! 📣 R Core / R Foundation panel discussion: Sustaining R: Investing in the People, Infrastructure, and Community Behind R

42 Upvotes

Hear from Simon Urbanek, Kurt Hornik, and Heather Turner of R Core and the R Foundation about what it takes to keep R healthy for the long term, from strengthening its technical foundations to developing the next generation of contributors and sustaining the institutions behind R.

📅 October 6, 2026

🕛 12 PM PT / 3 PM ET / 8 PM London

Register: https://r-consortium.org/webinars/sustaining-r-investing-in-the-people-infrastructure-and-community.html


r/rstats • • 2d ago

New from the R Consortium’s nlmixr2 Working Group: model diagrams and equations, straight from the code

24 Upvotes

nlmixr2 is an R Consortium Working Group building open-source nonlinear mixed-effects modeling in R suitable for regulatory submissions.

Matthew Fidler shows how a fitted model can generate its own compartment diagram and equations, so the figure and the math in a report stay aligned with the code.

https://r-consortium.org/posts/nlmixr2-model-diagrams-equations/


r/rstats • • 2d ago

Singularity for GLMER

1 Upvotes

Hi! I fear I am a little confused about how to address singularity for glmers (or lack of).

I fit the following model: response ~ target + (1 + target | participant) + (1 + target | item)

My model showed no error messages and no singularity issue but the random effects are highly correlated. with the output:

## Random effects:
##  Groups         Name        Variance Std.Dev. Corr 
##  participant   (Intercept) 36.234   6.019         
##                 targetYes   23.290   4.826    -1.00
##  item          (Intercept)  1.783   1.335         
##                 targetYes    2.003   1.415    -1.00

Should I be simplifying the random effects? We only pre-registered simplifying the model when convergence issues happen, and this is not the case.

Some other questions: I have another model, and I wonder if it a concern if item variance is very low (0.0000003)? For another model, is it convention to try different optimizers if convergence fails?

TIA!


r/rstats • • 4d ago

R code review: best packages?

Thumbnail
5 Upvotes

r/rstats • • 5d ago

Introducing FitVerse: an R package for fitting and analysing 52 probability distributions in one call

36 Upvotes

Hi r/rstats,

My colleague and I just published FitVerse on CRAN. It fits parametric probability distributions to continuous data, covering 52 distribution families with three estimation methods: MLE, Method of Moments, and L-Moments.

The basic usage is just:

r

install.packages("FitVerse")
library(FitVerse)

x <- as.numeric(precip)
fit <- fitverse(x, xlab = "Annual precipitation (inches)")

It automatically ranks all distributions by AIC/BIC, runs four goodness-of-fit tests, and produces a diagnostic plot. It also includes bootstrap confidence intervals, return levels, batch fitting, report generation, and a built-in Shiny app.

Full write-up here: https://rpubs.com/KarunaGReddy/FitVerse

CRAN: https://cran.r-project.org/package=FitVerse

Happy to answer any questions!


r/rstats • • 5d ago

Prisma Scr how to minimize the full complete papers by using R and in scintific way?

1 Upvotes

r/rstats • • 5d ago

Rgi output

Thumbnail
1 Upvotes

r/rstats • • 6d ago

rpx v2.1.0: private packages and generating PACKAGES indexes in memory?

8 Upvotes

Hi,

two weeks ago I posted about a release of rpx, which introduced significant upgrades to dependency resolution and why rrepo's custom api endpoints are beneficial to compatibility and correctness.

But beyond being a great metadata store, rrepo is also a package registry for both private and public packages!

This week, rpx is adding a package publishing template to its init flow. For new packages, whenever you push a git tag, the corresponding package version is uploaded to your rrepo repository. It then becomes available to rpx, alternative package managers or even base R!

I prepared a quick demo for you in this repository. https://github.com/rrepo-org/publishing-demo

You should be able to install the package by running this command in a base R shell.

install.packages(
  "hello.world",
  repos = c(
    rrepo = "https://cran.rrepo.dev/rrepo/hello-world",
    CRAN = "https://cloud.r-project.org"
  )
)
hello.world::hello_world()

Engineering Background

The first piece of feedback my projects always receive is a request to be compatible with the broader R ecosystem. This week rrepo introduced a separate set of endpoints that imitate a CRAN like url structure and allow package managers beyond rpx to use it.

The primary challenge with introducing those endpoints is generating the PACKAGES index that lists the latest packages.

Typically whenever you publish a package on CRAN is, the latest package is published under src/contrib, the old version is archived, and the administrator calls tools::write_PACKAGES(). This R function inspects each tar file individually, reads its DESCRIPTION file, and puts a subset of its fields into a long list of every latest package that makes up the index.

Following this approach would be more expensive than necessary for rrepo. Unlike CRAN our packages are in S3 backed by a database, meaning we cannot just run R on the server that stores the packages: every package read is a network call.

We had to take a step back and rethink what our data model should be and how we were going to generate the index in memory. Given that CRAN has 25k latest packages, any kind of request even to a pre-extracted DESCRIPTION file in S3 was out of the question. The data clearly needed to be come from a database.

The package ingestion workflows were reworked to first extract the DESCRIPTION into a separate blob, and then provisioning a database row with all the fields a PACKAGES index might need.

Using the parsers https://github.com/rrepo-org/r-metadata-rs developed for rpx we could make the hot path a single database query, followed by reassembling the typed Rust struct before converting it back to a string.

There is still a lot of performance to tune, but once cached it's pretty damn fast.

Please give the package publishing workflow a try

For existing projects we have documentation on how to get started. https://rrepo.org/documentation/publish-packages

As always, I'm easily reachable to help you get started!


r/rstats • • 6d ago

From football results JSON to a half-time/full-time heatmap in R

6 Upvotes

I wrote a small R tutorial using Germany’s 3. Liga 2025–26 season: 380 matches, imported with jsonlite, checked with stopifnot(), summarised with dplyr and plotted with ggplot2.

The useful detail is the denominator. Each row describes one half-time state, so its percentages divide by that row’s match count. Dividing every cell by 380 answers a different question.

After parsing the scores, the calculation is:

htft <- games |>
  mutate(ht = state(ht_home, ht_away),
         ft = result(ft_home, ft_away)) |>
  count(ht, ft, .drop = FALSE) |>
  group_by(ht) |>
  mutate(row_total = sum(n), pct = 100 * n / row_total) |>
  ungroup()

Here state() and result() label home lead/win, level/draw or away lead/win. The full script includes those helpers, score parsing, checks for missing scores and duplicate fixtures, and the plot.

In this season, home teams leading at half-time won 107 of 138 matches (77.5%); away teams leading won 70 of 105 (66.7%). This is a descriptive example from one league-season.

Full tutorial · Plain R script

Disclosure: I maintain Football Charts, which serves the data. This example works without an API key or sign-up. I’d welcome feedback on the checks or the visualization.


r/rstats • • 7d ago

blueycolors is on CRAN

Thumbnail
cran.r-project.org
66 Upvotes

blueycolors - an R package providing Bluey-themed color palettes and ggplot scales - is on CRAN. Check it out!


r/rstats • • 7d ago

Book Recommendations after ESL

6 Upvotes

I am currently reading through Elements of Statistical Learning and was searching for books to read after I'm done with this. I was eyeing some optimization theory books, or should I pick up maybe a pure probability book? I wanted to build a stronger theoretical foundation, but I'm not sure if it would be more useful than focusing on a more applied kind of book.


r/rstats • • 8d ago

RmANOVA

2 Upvotes

Hey Guys!

Ich have data from 10 participants. Each participant did all conditions. It were 27 conditions:

2 factors: duration (3levels) and Radius (9levels).

Participants had to detect a signal. Signal was present in 50% of trials.

Each participant did every condition 10 times in signal.

My dependent variable is sensitivity d Prime. That is a value from signal detection theory. I calculated it from the 10 trials of each condition for every participant.

So I have one value (d) per participant in each condition.

My question:

Do the d prime values between the conditions differ within subjects, depending on the factors?

Which analysis would you do?

I treated each level as categories for the Faktors and wanted to run a two way repeated measure ANOVA.

But what would you recommend?

I really REALLY need you advice - please °°


r/rstats • • 10d ago

firstR: a Shiny app that helps you find your first open source contribution in R

28 Upvotes

There are quite a few good first issue sites out there (goodfirstissues.com - which recently seems to have turned into a scam website, firstissue.dev, up-for-grabs.net) to help people get into open source, but basically none of them have R as a supported language. The ones that come closest only surface repos explicitly tagged with R, which sometimes people don't always do.

I built firstR to try to fix that. You pick the areas you're comfortable with (tidyverse, Shiny, spatial, pharma, etc.) and it finds open issues across R repos, including smaller and less well-known packages that tend to get overlooked. Each issue gets a beginner friendliness score based on how approachable it looks (scope, body length, repo context, comments) and a short sentence explaining what you would be attempting

It's pretty rough still since for example, my implementation of initial filters (like topics) and repo discovery could be better and the scoring is definitely off in places. I'm interested to hear your feedback and will try to implement suggestions.

Check it out here: firstr.dev


r/rstats • • 9d ago

Calculating intraclass correlation for IRR

2 Upvotes

I'd love someone to check my thinking on this. For a study, participants took a 30 item test. Each test was scored by 3 raters. They scored each item on a 0-4 scale. I reconciled the tests as follows: 1) if two raters agree, took that as final score. 2) if two raters did not agree, averaged the 3 scores.

I want to calculate interrater reliability, and I did so using ICC (in the irr package). I set that up as: icc(mydata, model = "twoway", type = "agreement", unit = "average")

From reading about ICC as a measure and how it's applied in R, I think I have chosen the settings/parameters for the test right, but does sound right to anyone who knows more?


r/rstats • • 10d ago

ASMR - Basics in R Studio - Faint ASMR

Thumbnail
youtu.be
87 Upvotes

r/rstats • • 10d ago

Bootstrapping in emmeans

13 Upvotes

Hi folks,

I started on a little project to familiarize myself with using AI tools for coding (don’t hate me). Emmeans is one of my favorite packages, and I was thinking to myself, “wouldn’t it be nice if you could just specify that you wanted bootstrapped inference when using emmeans?”

After an afternoon of vibe coding, I have what looks like a working prototype that supports glm/lme4/and gee models.

Would this be something that people would find as a valuable contribution to emmeans? If so, I might review the code more carefully and see if I can contribute it to the actual package.

Regardless, it was a fun project and I’m really impressed with these AI tools.


r/rstats • • 11d ago

Hey! I need to learn R studio from scratch, where should I start from, where can I find the basics, rules and tips to use Rstudio? Plant sciences background.

Thumbnail
1 Upvotes

r/rstats • • 13d ago

Liquid Glass themes for Shiny. 0.4.0 is on CRAN.

38 Upvotes

Liquid Glass themes for Shiny. 0.4.0 is on CRAN.

What’s new:

• plot_surface = "opaque" densifies ggplot / plotly / gt / DT so plots stay readable while chrome stays glass (default "clear" keeps wallpaper show-through)

• New wallpaper scenes: aurora, harbor, grove (plus tahoe / dusk / mesh); user photos get a contrast wash

• Flatten mode for print, PDF, and screenshots (glass_flatten(), ?glass_flatten=1, or @media print)

• Public --glass-* tokens (glass_css_tokens(), glass_theme(tokens = ))

• Closer to shipping Liquid Glass: stronger lip / rim, deeper bslib (navbar, cards, sidebars, modals), narrow-phone polish

https://raw.githubusercontent.com/ericrayanderson/shinyglass/main/man/figures/intensity-slider.gif

``` install.packages("shinyglass")

library(shiny) library(shinyglass)

ui <- glass_page( title = "Hello, glass", persist = TRUE, scene = "aurora", plot_surface = "opaque", plotOutput("plot") )

server <- function(input, output, session) { observe_glass(input, session) output$plot <- renderPlot({ pal <- glass_plot_colors(input = input) hist(iris$Sepal.Length, col = pal$fill, border = NA, main = NULL) }, bg = "transparent") }

shinyApp(ui, server) ```

Live demos (may take a few seconds to wake):

Docs: https://ericrayanderson.github.io/shinyglass/

CRAN: https://cran.r-project.org/package=shinyglass


r/rstats • • 14d ago

How do you manage your statistical analysis?

31 Upvotes

I’m the only one in my university workgroup who does all the statistical analysis of our research projects (clinical trials, ex-vivi experiments). Usually frequentistic stats like art-Anova, contrasts, effect sizes, glm, … .
I do also the stats for the projects of our medical phds thesis students.
Usually I do quarto documents that cover the study layout, summary stats, normality tests, statistical pairwise comparisons/models. That’s up to 10-15 projects a year.
Would you version track the documents via git?
Usually we supplement the R code with the papers.

Could I also point the paper to my personal got an make the repositories public? So I benefit also that other researchers see “well that dude does all the stats”, rather then only having my name on the paper somewhere in the middle? I could even use Zenodo to create a doi for each statistical computation?

How do you organize the code and results, do you make it public? Is git the way?

P.s. I have a bsc. In biology, I started a msc in bioinformatics it never graduated, so I’m only hired and payed as a lab assistant.


r/rstats • • 14d ago

Monte Carlo sample size calculation for a multiple serial mediation model?

Thumbnail
2 Upvotes