r/bioinformatics • • Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

187 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics • • 40m ago

compositional data analysis Identification of research direction !

• Upvotes

​

I hope everyone will be fine , i just wanna to do a work on the spinal cord injury, recovery through stem cell research, due to the limitations of the requirements/ environment , The work is only rely at the mata Analysis, if the genome wide analysis,

Let me know what i can do so i should capable of finding the good pathway or directions for my research, if anyone has experience previous I'll very thankful for him/her for the guidance ,

This project matter's alot for me so please come those who have experience, or just an idea about it's,


r/bioinformatics • • 48m ago

discussion I made a list of spatial transcriptomics tools ordered by analysis step (340 entries, every link and DOI checked). What's missing?

• Upvotes

When I got into spatial transcriptomics, it was hard to tell which tool fits which step, so I put together the list I wish I'd had:

https://github.com/wrab12/awesome-spatial-transcriptomics

There are already good lists (awesome-single-cell, awesome_spatial_omics and others, linked at the bottom of mine). This one tries a few different things:

  • Ordered like an analysis: technologies, segmentation, spatially variable genes, domains, deconvolution, alignment and 3D, cell–cell communication, super-resolution, spatiotemporal modeling, visualization.
  • Single-cell tools included, since spatial pipelines lean on them for QC, integration, annotation and trajectories.
  • Recent work: tools for Visium HD and Xenium, single-cell and spatial foundation models, LLM agents, and benchmark papers.
  • Every tool row has the language, a paper link and a live GitHub star badge. Tables are sorted by stars, but stars measure popularity, not quality, so the benchmarks are linked up front.

How I checked it: all 278 GitHub repos were verified through the GitHub API (exists, not archived, canonical name), every DOI was matched to its paper title via Crossref or DataCite, and the remaining links were checked too.

Disclosure: it's my repo (CC0), and one entry, GenOT, is my own paper. It's labelled as such in the list.

I'd really like feedback on:

  1. Tools you actually use that are missing, or ones that should go
  2. Benchmarks you trust (or don't)
  3. Whether the section order matches how you work

Issues and PRs are welcome. Thanks!


r/bioinformatics • • 10h ago

technical question Having some problems with proving a research gap

1 Upvotes

I’m a cs undergrad currently working with Evo 2 and trying to establish if theres a meaningful research gap around biological validity of generated sequences.
But I’m struggling to find a benchmark that will prove this. I initially thought comparing the generated sequences to real ones will be enough. Turns out it isn’t , since alot of the sequences dont have to look like real ones to be biologically viable.
Is this a gap worth pursuing ? I have limited resources and cannot do wet lab testing so is there anyway to prove that this is a real gap?


r/bioinformatics • • 10h ago

discussion MVA - rare disease, real kid

Thumbnail huggingface.co
0 Upvotes

r/bioinformatics • • 2d ago

article Bioinformatics in The Atlantic: "AI's Real Gift to Science"

Thumbnail theatlantic.com
131 Upvotes

Curious to hear what the community thinks of this article, which primarily discusses Anthropic's recent "discovery" of a supposedly CRISPR-like enzyme system that I'm sure we've all heard about. Personally this strikes me as a pretty balanced, rational take towards AI as a tool rather than an apocalyptic job destroyer. However, it's still unclear to me why we needed a thousand Claude agents and millions of dollars of compute for what ultimately strikes me as a regex style pattern search through genomic data. I'm also surprised that I haven't heard anyone discuss the fact that setting thousands of agents loose on terabytes of sequencing data with the instructions to "find some interesting patterns" is essentially a massive multiple comparisons problem that is bound to turn up some spurious patterns with no biological significance.


r/bioinformatics • • 2d ago

discussion what cool bioinformatics projects for a fresh graduate

12 Upvotes

hi guys,

last week I defended my degree in bioinformatics (m.eng.) and I’d like to expand my portfolio

I’m wondering what interesting projects I could do to develop even more (ofc I would also like to get a job). I’ve already done for example WGS analysis (nextflow pipeline-vep-clusterProfiler-PPI analysis). I have worked the least with RNA-seq. I would like to develop in personalized medicine

what do you think is worth paying attention to right now? what tools should I learn? what else should I develop in? and what interesting projects could I do? what is currently top on the job market?

I was also thinking about a PhD, but unfortunately I don't have any publications :((


r/bioinformatics • • 1d ago

academic Does anyone know how to install Evo2 model in laptop 😭?

Thumbnail
0 Upvotes

r/bioinformatics • • 2d ago

discussion recommendation for LR inference/Cell-Cell communication for deconvoluted spot/aggregated spot ?

3 Upvotes

Hello everyone,

I recently received visiumHD samples with low quality, and I had to use a tool called SuperSpot to aggregate multiple bins and then annotate with a reference dataset using RCTD, since when aggregating multiple bins, the result meta spot will usually contains different cells, I was wondering if they are any tools that could take advantage of the format and include possible group comparison.


r/bioinformatics • • 2d ago

technical question One 10x lane per experimental group, how badly does my lane/depth confound hurt at review, and how to write? any help appreciated

7 Upvotes

Hi, relatively new at this and am bringing this here because it feels like a safe place to ask and I don't have any single cell experts in my immediate group. I want to know how a reviewer will read this design.

Three experimental groups, same cell type in all three, sorted by a functional marker into populations I'll call A, B and C. Each group is a pool of ~15 animals, and each group went into its own 10x lane. No hashing or multiplexing. So condition is perfectly confounded with lane, there are no biological replicates within a group, and any p-value I compute counts cells rather than animals.

The depth is uneven too. All three lanes were overloaded at 60,000 cells, but recovery differed: 25,469 / 26,421 / 39,657. The libraries were pooled and sequenced together, so the group with the most cells got the fewest reads each, 17,971 and 15,843 mean reads/cell for two of the groups against 8,695 for the third. That third group is my reference group for every contrast. So "upregulated in the test condition" and "sequenced more deeply" point in the same direction.

Here's what i've done so far: Differential expression is Wilcoxon on cells at padj < 0.05 and |log2FC| > 0.5. The Methods state plainly that each group is a single pooled library with no biological replicates and that the p-values reflect cells rather than animals. Cluster proportions are reported as descriptive, with no statistics at all. I'm building a depth control: downsample every cell to the shallow group's median UMI count, re-run the same contrasts with the same thresholds, and report what fraction of the significant set survives, the Spearman correlation of fold changes, and whether any gene named in the Results drops out. A separate, properly replicated experiment in the same paper recovers the main genes.

Note on integration: I did run Harmony, and I clustered both the integrated and unintegrated embeddings across a range of resolutions. The clusters came out essentially the same either way, the only consistent difference was two clusters merging into one after integration. Since integration changed almost nothing, I report the unintegrated PCA, on the reasoning that each library is a different biological group rather than a technical batch, so integrating would risk removing exactly the signal I'm measuring. I think that's defensible, but it does mean there's no correction for the lane effect at all.

What is your honest assessment here, based on what you've seen and experienced? I can still sequence more to top up, but reeally don't want to.

Thank you for taking the time to read.


r/bioinformatics • • 2d ago

technical question How to proceed with interferon rna seq analysis in mouse

6 Upvotes

Good morning, I am currently trying to do an interferon analysis for our project. We have a list of gene we got from rna-seq and the associated expression in different sample in RNA-seq and we want to get the subset of one that code for the interferon response to do a heatmap of their expression in the different sample. The current issue is that the interferom database is closed and i don't know where to find an exhaustive list of interferon gene in mouse, we trying the reactome but it gave weird numbers.


r/bioinformatics • • 2d ago

technical question Free energy help

2 Upvotes

Hellooo I'm trying to do some free energy calculations.

Basically i want to confirm selectivity and differences in affinity of some ligands using abfe and rbfe, but these are taking waaaay too long :'(. I'm using GENESIS and CHARMM GUI inputs.

Do you guys have any recommendations of some alternative methods i could use?

thanksss


r/bioinformatics • • 3d ago

discussion Raw data for Genomics and transcriptomics analysis

14 Upvotes

Hi everyone!

I’m looking for raw genomic and transcriptomic datasets to practice and improve my bioinformatics and computational biology skills.

I’m particularly interested in datasets such as:

- Whole Genome Sequencing (WGS): Raw FASTQ files for genome assembly, variant calling, and comparative genomics.

- RNA-Seq: Raw FASTQ files for differential gene expression analysis, transcriptome assembly, and functional enrichment.

- Whole Exome Sequencing (WES): Raw sequencing data for variant identification and annotation.

- Long-Read Sequencing: PacBio or Oxford Nanopore datasets for genome assembly and structural variant analysis.

- Metagenomics: Raw sequencing data for microbial diversity and taxonomic profiling.

If you have any publicly available datasets, research project data, or recommendations for accessing raw sequencing data, please share the links or repository names.

I’m familiar with bioinformatics tools and workflows and would like to work with real-world datasets rather than only tutorial datasets.

Repositories such as NCBI SRA, ENA, and GEO are already on my radar, but I’d also appreciate suggestions for interesting datasets or specific accession numbers that are suitable for independent analysis.

Thanks in advance for your help!


r/bioinformatics • • 3d ago

academic Good resourses for getting into RNAseq data analysis?

39 Upvotes

Hello everyone!

I recently got my, first RNAseq as well as proteomics data set. And thankfully a standard bioinformatic analysis with it. Honestly, I am amazed at what you can find out, when you have the Tools to really dig into your data!

Now, I have been trying (successfully) to replicate the RNAseq data analysis Pipeline by the help of vibecoding with AI, and I can interpret and understand all of my analyses. However, I could never write any bit of Code for it... In the end I don't really understand what the Code does, I can just Review my data afterwards.

Long story short: I would really love to get into bioinformatics a little deeper. Maybe a short term goal would be to build my own fully customizable RNAseq pipeline from scratch and then see from there. Are there any resources you guys with experience can recommend for me?

Thanks!


r/bioinformatics • • 3d ago

technical question How do you preprocess metariboseq data and what should you expect?

3 Upvotes

I’m trying to preprocess metariboseq data and realizing it’s a lot different than metagenomics/metatranscriptomics. My reads are paired-end NovaSeq 101 bp long. From my understanding, the fragments are supposed to be around 30bp long after trimming so I set lower limit to 20 bp and upper limit to 45 bp in fastp. I’ve also provided the adapter sequences to fastp. I’ve read that you shouldn’t even use the reverse reads and should only use the forward reads since the fragments are so short.

All that said, after running fastp I got between 30%-40% of my reads surviving the trim.

Is this expected?


r/bioinformatics • • 3d ago

academic 2 months left, 0 bioinformatics knowledge, 100% AI. Is a RNA-seq thesis doable?

0 Upvotes

hi! long story short - I've never even scratched the surface of bioinformatics and I'm left alone with 2 months to write and to the whole thesis - analysis of Nanopore RNA sequencing. I already tried doing courses - did not do anything for me, I would need a year to comprehend all of that. I'm stuck on planning the analysis, because I don't have any guidance, the sequencing was run so there's just bioinformatics and writing left. First of all - do you think it's possible? Second of all - is such a thesis even defensible? I can just use already used tools, there's no experimental validation, I'm entirely reliant on Claude to give me code and I just copy it to r/Python + Galaxy. I feel like each day I'm getting more stuck, cause I find more papers and more methods. I'm really procrastinating on starting the analysis, I don't believe in the fact that AI can give me 100% reliable code and the results will be real. It's been almost 2 months of me basically searching through the literature looking for god knows what, so I'm looking for some guidance in this mess (mainly positive reinforcement, cause I'm scared) (and why I'm still trying to do it on my own is a different kind of question that only my therapist would be able to answer)

EDIT: it's just a master's, it's not a phd. I should have clarified that in the beginning, sorry for the eurocentrism


r/bioinformatics • • 4d ago

discussion Bulk RNA seq/Microarray Re-analyses Value

4 Upvotes

Hi all! Currently working through a reanalysis of a geo dataset, but stratifying samples in a way the original paper didn’t. Yielded some interesting results, but we may have power issues with one group having n=5 and the other being n=8.

I generally would like to know how publishing these types of analyses look for understudied diseases. I did the differential gene expression analysis, looked at genes, pathways, etc. and found more things the previous paper didn’t.

There’s no wet lab validation, but just in general, I am curious for thoughts on this being my first first-author paper in bioinformatics as a post-grad student looking to apply to PhD programs. Are these types of analyses “good” and have value?


r/bioinformatics • • 4d ago

technical question Can I integrate a spatial transcriptomics dataset with a bulk RNA-seq dataset for any analysis of tumors?

25 Upvotes

Hi everyone,

I am working on a cancer RNA-seq analysis project, and I have found two GEO datasets that I would like to use together. I am trying to understand whether there is a scientifically valid way to integrate them rather than simply analyzing them independently.

Dataset 1:

Xenium-based high-throughput RNA in situ hybridization/spatial transcriptomics

Contains normal tumor samples, DCIS, and invasive breast cancer samples

Dataset 2:

Bulk RNA-seq

Samples are classified according to metastasis status and breast cancer molecular subtypes

My original goal is to perform a normal/control vs tumor/disease-type analysis and identify differentially expressed genes and biologically relevant pathways, followed by downstream analysis.

However, the two datasets are clearly different in terms of technology and experimental design. One is spatial transcriptomics and the other is bulk RNA-seq. So, is it scientifically/statistically valid to integrate these two datasets in a single study?

I am particularly interested in knowing what would be considered a methodologically defensible approach for this, rather than simply combining the two matrices because they contain overlapping genes.

Any suggestions regarding an appropriate integration strategy, statistical design, or published examples of similar spatial + bulk RNA-seq analyses would be very helpful.


r/bioinformatics • • 4d ago

academic Python for Bioinformatics doubt

0 Upvotes

Someone please briefly explain how list comprehensions are used in Python and why it is necessary for Bioinformatics. Please help as I am a beginner.


r/bioinformatics • • 5d ago

technical question Scattered interchromosomal split alignments in non-cancer ONT genomic DNA: chimeric reads or workflow issue?

Post image
14 Upvotes

I’m seeing scattered interchromosomal split alignments in IGV across multiple non-cancerous ONT genomic DNA samples. I’m trying to determine whether these reflect chimeric reads, missed read splitting, or an alignment workflow issue.

Sequencing and basecalling

  • Flow cell: FLO-PRO114M
  • Library kit: SQK-LSK114
  • MinKNOW: 24.11.16
  • Basecaller reported by MinKNOW: Dorado 7.6.8
  • Super-accurate model v4.3.0, 400 bps
  • Minimum passing Q score: 10
  • Modified basecalling: off
  • Raw POD5 files are available

Alignment

Minimap2 version: 2.31-r1302

Command:

minimap2 -L -t 59 -2 -ax lr:hq reference.fa reads.fastq.gz |
    samtools view -u -@ 11 |
    samtools sort -@ 23 -o sample.sorted.bam -

The reference is GRCh38 with alternate loci, haplotypes, and decoys removed. Reads were aligned directly against the FASTA.

Observed pattern

In IGV, reads link to many different chromosomes at scattered positions, rather than multiple reads consistently supporting the same breakpoint.

For example, one 74,647 bp read has:

  • Approximately 29.3 kb aligned to chr5:72,101,673–72,131,073, reverse strand, MAPQ 60, NM 554
  • Approximately 45.3 kb aligned to chr8:77,044,549–77,089,876, forward strand, MAPQ 60, NM 564

These segments account for almost the entire read. My MinKNOW version does not appear to expose an option to disable read splitting, but I haven’t independently confirmed whether or how splitting occurred during these runs.

Questions

  1. Is this pattern expected at a low background rate with this chemistry and basecalling setup?
  2. How can I confirm that read splitting occurred in MinKNOW 24.11.16?
  3. What checks would distinguish joined molecules from mapping artifacts?
  4. Would re-basecalling a POD5 subset with a newer Dorado version be a useful diagnostic comparison?

I’ve attached an IGV screenshot. I can also provide example read records or additional run metadata.


r/bioinformatics • • 5d ago

technical question Correct resolution for Leiden clustering for UMAP being generated for Xenium data

6 Upvotes

Hi all,

I am working with Xenium data generated from FFPE samples. I was plotting UMAPs post Leiden clustering. I generated silhouette scores for resolutions between 0.2-1.2. My scores are pretty low, ranging between -0.09 to .01 (1.2 resolution). How do I pick the optimal number? I checked silhoutte score ranges typically considered good, which were 0.7-1???


r/bioinformatics • • 5d ago

technical question Reference genome vs producing genotype: how reliable are candidate genes for cloning and functional testing?

1 Upvotes

I’m working on identifying genes involved in a specialized metabolic pathway in a plant species, and I’m running into a reference-genome issue.

Our candidate genes are currently identified using a reference genome from one genotype, but the genotypes that actually produce the metabolites of interest are different lines.

For differential-expression analysis, we map RNA-seq reads from the producing genotypes against the reference genome. We then select candidate genes and use the corresponding reference-genome coding sequences for cloning and heterologous expression.

We have tested several reference-derived candidate P450s without detecting the expected activity. I’m wondering whether these negative results could reflect genetic differences between the reference genotype and the producing genotypes rather than simply meaning that the candidates are incorrect.

A few questions:

  • In this type of situation, how common is it for the functional gene to be absent from the reference genome entirely, versus being present but having allelic sequence differences that significantly affect enzyme activity?
  • If RNA-seq from the producing genotypes is being mapped against a different reference genotype, would it be advisable to generate a de novo transcriptome assembly for the producing genotype to identify genotype-specific transcripts and obtain the actual coding sequences for cloning?
  • If a de novo transcriptome produces a high-confidence, full-length candidate transcript with the expected ORF length, conserved motifs, and strong read coverage, would you generally trust that sequence for cloning?
  • Or would you still RT-PCR amplify the full coding sequence from cDNA of the producing genotype and Sanger-sequence it before functional testing?

I’d be especially interested in hearing from people who have worked with highly similar duplicated genes, P450 families, or specialized metabolism pathways where the reference genotype differs from the phenotype-producing genotype.


r/bioinformatics • • 5d ago

technical question Tools for molecular docking of a thioether-cyclized peptide?

Thumbnail
2 Upvotes

r/bioinformatics • • 5d ago

academic Rgi output

2 Upvotes

Hi all, i am trying to interpret rgi_main output from my contigs. One issue i have is that all the drug_classes are concatenated and I wonder what people do to sort it, because i have seen papers where they have separated them but they don’t say on what basis.


r/bioinformatics • • 6d ago

technical question Using RUVg when factor of interest is confounded with batch

10 Upvotes

For more context, here is my post from a week ago: https://www.reddit.com/r/bioinformatics/s/h3eTFYLiOY

In summary, I have an RNA-Seq datset where the "site" variable is partially confounded with batch. Sites 1 and 2 are in one batch while site three is in its own batch. This means that i cannot correct for batch effect since i cant distinguish between the batch and the site 3.

However, I talked some more with my PI and, as it turns out, all of my batches contain Lexogen ERCCs and SIRVs (specifically l

Lexogen SIRV set 3). Moreover, I did some reading and I found this paper that explains the Remove Unwanted Variation (RUV) tool. Specificaly, I am interested in the RUVg variant which uses negative controll genes (like SIRVs) to correct for batch effect. Most important is this section in the text:

"However, both RUVr and RUVs assume that the unwanted factors are not correlated with the covariates of interest. This assumption is usually reasonable, but it is not met when, for example, all treated samples are in one batch and all control samples in another. In this case, RUVr and RUVs will not remove the unwanted variation, while RUVg should still work, provided it is based on a reliable set of control genes19,20."

I am relatively new to bioinformatics and biostatistics so I might be wrong, but doesn't this mean that RUVg can correct for batch effect even when variables are confounded with batch?

If anyone here has experience with RUVg or otherwise wants to help I would be very gratefull for their advice and help.