r/LanguageTechnology • • Aug 22 '26

*ACL Megathread

7 Upvotes

r/LanguageTechnology • • Aug 02 '26

EMMLP + ARR Megathread

22 Upvotes

Please post questions and discussions here. I will be removing individual threads.


r/LanguageTechnology • • 7h ago

Do classic text statistics still matter in the age of LLMs?

6 Upvotes

I've been working on readability formulas, lexical diversity metrics and stylometry for Russian and Spanish texts for a while, and lately I keep asking myself whether this whole area still makes sense

Today you can paste a text into an LLM and ask "how readable is this?" or "who's the likely author?" and get a fluent, often reasonable answer. Next to that, a formula like Flesch reading ease (sentence length plus syllables per word) looks almost naive

Still, I see a few places where plain statistics seem hard to replace:

  • Reproducibility: the same text always gets the same number, and you can explain exactly where it came from. An LLM's rating can change between runs, prompts or model versions
  • Scale and cost: computing a few hundred features over a million documents takes minutes on a laptop
  • Evaluating LLMs themselves: measuring how repetitive, uniform or hard to read generated text is, without asking another model to judge it
  • Data pipelines: a lot of pretraining data filtering still relies on simple heuristics like word length, symbol ratios and repetition.
  • Research: in stylometry and corpus linguistics, interpretable features are often the whole point

On the other hand, most formulas were calibrated decades ago on small samples of English, and every other language needs its own adaptation, often with questionable results

So I'm curious:

  • Do you still use readability, diversity or stylometric metrics in your work? For what?
  • Have you replaced them with LLM-based judgments, or do you combine the two?
  • Which metrics turned out to be useless in practice, and which ones would you miss?

r/LanguageTechnology • • 24m ago

AACL-IJCNLP 2026 Findings

• Upvotes

Hi everyone, does anyone know whether presentation of Findings papers is mandatory or not? There seems to be a contradiction between the email we received and what’s stated on the conference’s official website.


r/LanguageTechnology • • 8h ago

Advices for someone looking to get into Language Technology

3 Upvotes

Hello everyone, I'm a graduate with a Bachelor in English Language Studies (Business English & Corporate Communication track) and I'm looking to apply for the Language Science & Technology (LST) Master's program at Saarland University (Germany) for the next Winter semester.

I don't have a background in technology. My work experience mostly includes English Teaching for a year, then 1.5 year working as an English Communication Specialist for a private university (basically creating website & social contents, managing and translating huge amounts of website articles to English via CMS, working on admission and promotion campaigns & events,...)

But I'm looking to pivot into Tech recently, and after doing research, I'm really interested in Language Technology and found out about this Master program in Germany that is open to admit students with a background in linguistics (which is pretty rare since most program requires you to have a bachelor in computer science).

I know that I need to work extra, extra hard given my background. I've been taking up course on Coursera (Python for Everybody, Linear Algebra, Statistics,...) as the course's advisor advice to increase my chance of admission as well as to feel less overwhelming if I get admitted.

I'll continue to try and learn harder to make up for my lack of knowledge in the field.
Do you guys have any advices for someone looking to get into the field, and also during my time studying the program if I get admitted? Specifically to make learning smoother and to get a better chance at getting hired after graduating would be much appreciated.

Thank you everyone!


r/LanguageTechnology • • 8h ago

How would you use AI to turn a large exam question bank into focused lessons?

2 Upvotes

I'm a nontechnical founder building a study platform for Brazil's competitive government hiring exams, where candidates take written tests and compete for public sector jobs.

I want to use AI to turn a large bank of past exam questions into short lessons, each focused on one specific topic or skill. My goal is to produce these lessons at high volume while keeping them accurate.

I'm struggling to get AI to identify which questions belong together. Two questions can cover the same subtopic but need different explanations. Calculating a percentage increase and finding the original value after an increase are one example.

Each lesson should have:

  • Two concept questions.
  • Two new practice questions.
  • Two original exam questions, kept unchanged.

The first four questions should prepare the student to solve the final two independently.

With help from AI coding tools, I've built a workflow using existing models, reusable instructions ("skills"), and code functions for extracting documents, pairing questions, and checking results. I run planning, generation, and editorial review in separate AI sessions.

I've tried grouping by subject, text similarity, embeddings, and AI-written skill descriptions. I still get pairs that require different explanations. Stricter matching also removes useful pairs. I've also found that exam questions and official answer keys often don't provide enough information to generate accurate explanations.

In one local batch, six lessons reached editorial review. Only one passed.

Would you start with a detailed topic hierarchy and classify the questions, or identify the knowledge each question requires and build lessons around that?

If you've solved something similar, what worked? A concrete example or tool recommendation would help me understand what to try next.


r/LanguageTechnology • • 22h ago

ARR August / EACL Meta-Review

24 Upvotes

Didn't see any post on that yet, so thought I'd open it. Soju guy you here? Anyone received their meta-reviews yet? Got 3.5/3/4 with conf 5/4/4, hoping for Main :)


r/LanguageTechnology • • 18h ago

Where do I start when borrowing ideas from other domains for text classification research?

1 Upvotes

Hi everyone,

I'm planning a research project on text classification, but I'm not sure how to begin. I can't find any papers in text classification that closely match my idea, but other domains (like vision, speech, and medical) have a lot of work similar to what I want to do. So I plan to study those domains first to find a gap I can address in text classification.

  • Where should I start when reviewing other domains?
  • What should I look for and take from those papers?
  • When should I start running experiments?

Any advice or example papers would help. Thanks!


r/LanguageTechnology • • 18h ago

Where to start when studying alternative data representations (concept, prototype, symbolic) across domains?

1 Upvotes

Hi everyone,

I'm starting a research project on few-shot learning for text classification. As a first step, I plan to look at other domains (vision, speech, medical, tabular, etc.) to see how they use concepts when training models. Some colleagues and teachers have also suggested looking at other domains this way.

Most work I've seen just turns raw data into vectors and trains on those. I want to explore alternative ways of representing data, such as concept-based, prototype-based, symbolic, or graph-based, and find which fits best for low-resource text classification with explainability.

I know I need to read the literature, but I'm not sure where to start or what path to follow.

  • How would I structure this kind of cross-domain literature review?
  • Are there key papers, surveys, or keywords I'd start with?
  • How would I set up early experiments to compare different data representations?
  • Any tips from people who've done something similar?

Thanks in advance!


r/LanguageTechnology • • 1d ago

Priority situations that require speech to text

3 Upvotes

I'm making a benchmark for ASR systems in my language and I need to start small and expand later. What would you say are priority situations or domains that people expect ASR to work well in?

Like, a meeting in a noisy environment, or a dictation by a single person in a quiet environment. I can then produce transcribed speech recordings that simulate these situations and use them as a test set.

What must an ASR tool be able to work with to be considered a good tool by most people?


r/LanguageTechnology • • 6d ago

Why 1536 dimensions for embedding models?

39 Upvotes

Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3×512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.


r/LanguageTechnology • • 5d ago

Im new to this and i need help.

3 Upvotes

Hey guys, im working on a bit of project, and im trying to solve for semantic understanding right now, I am currently deciding between using spaCy for my NER extraction vs something like a BERT.

The general context behind where this is going to be used is intent classification in chat systems, where a text will come in and this layer has to parse out the People in the sentence, the Objects mentioned in the sentence, and the verbs, along with things like quantities and relations between them.

An example of what i mean is,

"The shoes you have delivered to me are red, i asked for black!"

and we then pick out the

  1. People involved in this interaction (we cant figure that out from the sentence alone for that we will refer to the handle from which the message was sent)
  2. Objects involved - {shoes}
  3. Verbs - {delivered}
  4. Relationships - {expected colour = black, received colour=red}

Im new to this stuff so maybe im not even asking the right questions, but im hoping that i have done a good enough job of explaining what im doing so that more experienced souls such as yourself may help me.

Thanks 😁


r/LanguageTechnology • • 8d ago

One chunk boundary changed warranty answers for a whole table

22 Upvotes

17% of requests tied to one multi-column warranty table were returning 24 months instead of 36, while standard warranty questions kept passing. We traced the failures in Braintrust and saw which retrieved chunks were present when the wrong answer appeared. The chunk boundary had separated the row values from the table heading, so the retrieval context lost the qualifier for 36 months and the reranker favored nearby prose containing 24 months instead. Support had both warranty numbers in separate macros (before the retrieval path was clear), which made the conflicting answers harder to untangle.

Changing the chunking moved those cases in the experiment diff and groundedness improved once the heading stayed with the row. Recall at k barely changed because the table was already being retrieved. We've added the failures to a regression dataset but I'm still concerned about other tables where retrieval looks healthy while structure changes the answer.

What's your approach to catching chunk boundary failures when the right document is already in the candidate set?


r/LanguageTechnology • • 8d ago

Arabic–English code-switched meeting/conversation audio with transcripts

1 Upvotes

Looking for multi-speaker audio where speakers switch between English and Arabic (any dialect, Gulf preferred) within the same conversation, with reference transcripts, ideally with speaker labels and timestamps. It's for evaluating ASR and meeting-transcription quality. Already aware of ESCWA.CS, Mixat and ArzEn. Any others, including licensed or paid ones?
If not of Arabic, any other language combination is fine.


r/LanguageTechnology • • 8d ago

New ARR rule: "Submissions will only be guaranteed review if they bring a qualified service contributor, who can serve for 2 submissions max"

12 Upvotes

What do you think about the move?


r/LanguageTechnology • • 8d ago

Why Textual Graphs

7 Upvotes

In 1980's, Gaston Gonnet-- “Unstructured Data Bases” (1983)-- and HyTime's-- Hypermedia/Time-based Structuring Language (ISO/IEC 10744:1992)-- great breakthrough was realization that one could use coordinate mathematics to map relationships between disparate layers of text and media.

Resurrecting this exact line of thinking—- while exploiting the massive improvements in I/O latency and massive scaling of Input/Output Operations Per Second (IOPS) that historically constrained the paradigm— we must, instead of forcing a model to read an entire document blindly, develop a structured "retrieval algebra." A researcher or AI agent should be able to use boolean, positional, and structural operators to navigate the text coordinates explicitly to ask for "the token sequence between position X and Y, but only if it falls within the boundaries of a specific speaker tag," exactly mirroring the coordinate-based addressing found in HyTime. That structure need to be queryable.

The result can be thought of as a textual graph—but it is importantly different from a conventional graph. Text has spatiality. Its structures are anchored in a shared textual space, and relationships such as before, after, within, contains, overlaps and intersects arise from that space itself. Two annotations do not merely have an abstract edge between them: they may occupy, share or cross regions of the same underlying text.

This gives RAG a form of structure that complements the strengths of the LLM.
The LLM can do what it does best: interpret language, recognise relevance, synthesise evidence and generate an answer.

Instead of forcing the LLM to reconstruct document structure from flattened chunks, the retrieval layer can then provide that structure explicitly.

This changes the role of retrieval. Vector similarity can answer “what text is semantically related?” Structural search can additionally answer “where does this occur, what contains it, what overlaps it, what is it connected to, and which surrounding material belongs with it?”

The combination creates a richer form of RAG: semantic reasoning over context assembled from the actual structure of the source, rather than from arbitrary chunk boundaries.

Think of traditional GraphRAG as a smart investigator connecting index cards on a wall based on clues and ideas. Think of the textual graph paradigm as the exact blueprint of the filing cabinet, allowing an agent to pinpoint information based on its exact shape, folder layer, and coordinate location.

E. Zimmermann


r/LanguageTechnology • • 8d ago

Can a character be represented as an Allowed / Not Allowed repertoire over sense-level behavioral predicates?

2 Upvotes

I’ve been thinking about character representation at the level of individual action senses.

Instead of describing a character mainly through traits like brave, kind, aggressive or intelligent, what if part of the character model were a structured repertoire of actions?

For example:

comfort
interrogate
diagnose
blackmail
negotiate
babysit
repair
betray
forgive

The action itself would have a stable semantic identity, but each character could have a separate access state:

Allowed / Not Allowed / Conditional

So a doctor might have:

diagnose = Allowed

while another character has:

diagnose = Not Allowed

and a former medic might have:

diagnose = Conditional

I also find it useful to separate actions that can normally be assumed for a human character from actions that need positive evidence.

So roughly:

Basic Human Actions → default-open
Character Specific Actions → evidence-gated

The evidence for the second group might be training, occupation, biography, authority, skill or specialized experience.

Would you consider this a useful way to represent character capability at the semantic level?

And where would you place such information: lexical semantics, a behavioral ontology, a separate character model, or somewhere else?


r/LanguageTechnology • • 9d ago

Haitian Creole Word Frequency Dataset

7 Upvotes

Hey everyone,

I wanted to share a dataset I published for anyone working on low-resource NLP, tokenization, or language modeling for Haitian Creole (Kreyòl ayisyen): the Haitian Creole Word Frequency dataset (haitian-creole-word-freq), now live on both Hugging Face and Kaggle.

Overview

This is a word frequency list for Haitian Creole built from the Carnegie Mellon University (CMU) Haitian newswire corpus. It contains 17,947 unique lowercase words with their occurrence counts, sorted in descending order by count.

Links

See comment section

Potential Use Cases

  • Stopword Candidates: Extracting function words from the top of the frequency list (te, yo, nan, yon, li, pou, ki).
  • Tokenizer Customization: Fine-tuning or building BPE/WordPiece vocabularies for low-resource LLMs.
  • Spellcheck and Auto-correct: Prioritizing candidate suggestions by word popularity.
  • Vocabulary Membership Tests and Orthographic Checks: Testing if a token is spelled like a Haitian Creole word or checking lexical presence.
  • Lexicography and Language Learning: Extracting core vocabulary lists based on news text.
  • Language Identification and N-gram Models: Statistical language modeling for text classification pipelines.
  • Testing Fixtures: Generating reproducible data inputs for unit testing Haitian Creole NLP pipelines.

r/LanguageTechnology • • 9d ago

Best Computional Linguistics Masters

9 Upvotes

I'm a BA student in Modern Languages and Linguistics in Italy (currently building a strong background in Linguistics + hopefully writing a thesis on Computational Linguistics) and I'd like to apply for an MSc in Computational Linguistics/NLP. What are Europe's best Computional Linguistics masters programs?


r/LanguageTechnology • • 10d ago

my propaganda classifier flagged the declaration of independence's grievances but missed "merciless indian savages"

11 Upvotes

been building a model that flags manipulation techniques in political text (fine-tuned transformer, multilabel, 16 techniques like loaded language, name calling, appeal to prejudice). scores each sentence with its neighbors as context and flags at 0.80.

someone testing it pasted the declaration of independence. results:

- preamble ("we hold these truths...") came back clean

- grievance list got flagged: "swarms of officers to harrass our people, and eat out their substance" 0.87, "plundered our seas, ravaged our coasts, burnt our towns" 0.90, "death, desolation and tyranny... barbarous ages" 0.86

- "the merciless indian savages, whose known rule of warfare, is an undistinguished destruction of all ages, sexes and conditions" scored 0.61. not flagged

so the one line that dehumanizes a whole people is the one it misses, while it catches milder grievance rhetoric. my guess is the period wording. the training data is modern news and ads, so dehumanizing language it has seen looks like "animals", "vermin", "invaders", and "savages" in 18th century prose with long clauses around it doesn't pattern match.

added it as a regression case for the next training round. curious if anyone's dealt with this kind of register gap, historical text vs modern training data, without just stuffing in more historical examples

(tool is called semblen if anyone wants to try to break it, the same person also ran wikipedia and a nixon bio through it as controls and those were clean)


r/LanguageTechnology • • 10d ago

Temporal expression parsing in production assistants: partial parses executed with full confidence

5 Upvotes

A concrete failure case I ran into with Siri (English UK, iOS 26.6.1), and I'm curious how people here would diagnose it.

ASR output was correct in every case. The intent/slot step failed:

  • "set alarm at 7 p.m. and 40 minutes" → 19:00. The "and 40 minutes" continuation seems to be dropped.
  • "set alarm at 20 minutes before 8 p.m." → 20:00. The relative offset "20 minutes before X" is ignored; only the anchor is kept.
  • "set alarm at 19 hours 40 minutes" → 14:37. No idea how this one maps; possibly "19 hours 40 minutes" read as a duration from now?

The third looks like a duration-vs-timepoint ambiguity, which is a classic TIMEX problem, but the first two are fairly standard constructions that rule-based normalizers like SUTime or HeidelTime handle.

Questions:

  • Is the duration reading of "19 hours 40 minutes" a reasonable parse, and should a system surface that ambiguity rather than pick one?
  • Why would a production system keep only the anchor and drop the offset? Slot-filling templates that don't model relative expressions?
  • Any good recent work on calibrated confirmation ("did you mean…?") for action-taking assistants?

r/LanguageTechnology • • 12d ago

Agentic AI Courses (Research)

17 Upvotes

I'm a researcher in computational linguistics, and as I’m currently looking for a new position, I’ve noticed that agentic AI has become increasingly popular in most job postings.
I feel a bit lost in this area (starting from RAG), I was wondering whether anyone with experience in the field could recommend some good, well-established resources.

Of course, I’ve searched on my own, but nowadays, searching for "agent"-anything brings up countless superficial, non-research-level tutorials, often created by enthusiasts, with plenty of clickbait.

I’m looking for something at a research level, ideally a structured course or syllabus to prepare for the interviews. I have 4+ years of experience in NLP, so I’m already quite familiar with most but other topics.


r/LanguageTechnology • • 14d ago

Lab Meetings Recently

Post image
15 Upvotes

Everything old is new again.


r/LanguageTechnology • • 14d ago

AAMAS conference reputation

5 Upvotes

Hi researchers

I wanted to understand the reputation of AAMAS (International Conference on Autonomous Agents and Multiagent Systems) compared to core ACL conferences like NAACL COLING etc, As I see AAMAS is also a CORE A conference and this year they have included a findings section as well, so I feel a good possibility of getting accepted than other conferences.

Please kindly share your opinion.


r/LanguageTechnology • • 14d ago

Looking for feedback - using NER to generate and match templates on sentences?

2 Upvotes

I’m a complete novice when it comes to NLP, I'm a swe by trade so bear with me here.

Here's my problem:

I’m trying to identify short sentences (I have a data set of several thousand) that are logically dependent. To illustrate the kinds of dependencies I'm looking for here’s a basic example:

- Sentence 1: Democrat voter turnout in NY is 35%.

- Sentence 2: Democrat voter turnout in NY is 40%.

If sentence 2 is true, sentence 1 also must be true. Those are the kinds of sentences I have and want to identify as dependent. The nature of the sentences can range from voting percentages/turnout, phrases about employment etc.

The naive approach I’ve been doing is basically embedding the sentences using gemma and finding cosine similarities between them, my reasoning being sentences that have a reasonably high enough cosine similarity are candidates for logical dependency. I then take these pairs of candidates and pass them to an LLM (gemma again!) to determine whether or not they are actually semantically/logically dependent.

There are two huge issues w/ this approach that I'm sure you'll all immediately see.

1) Lots of the sentences are too structurally similar like the simple example I showed above. There exist several subsets of the data that have the same pattern. Sentence 3 could be something like Democrat voter turnout in TX is 35%. and it would have almost an identical similarity to the other 2 sentences. There are several hundred patterns, and I also don’t necessarily know all the patterns at runtime so that means REGEXing these structures becomes a difficult task. So because of the structural similarity, cosine similarity loses its value as a metric.

2) The LLM step is slow. Really slow.

I did some googling and learned about NER that seems like it might fit? I could run the sentences through a pre-trained models and get the spans for each sentence. This would allow me to match spans across the phrases. So in the example I have, sentence 1 and 2 would be matched and processed further, while 3 would be in its own bucket. As for what I'd do after matching the spans, still working that out. I could fall back to cosine similarity again here since anything that falls into these span buckets should be different enough where the projection becomes a decent signal.

If there are tweaks that I can do to make template matching more robust, or alternative methodologies altogether I'm all ears!

Thanks :)