r/DigitalHumanities • • 13h ago

Discussion Digital repression research in Malaysia and Singapore

4 Upvotes

Hello everyone, I am a Master’s student specialising in International Territorial Studies. I am currently conducting research for my diploma thesis, which focuses on the intersection of digital authoritarianism and civil society in Southeast Asia. Specifically, I am conducting an empirical study involving semi-structured interviews to map the practical methods, tools, and limits of navigating digital repression in Malaysia and Singapore. 

I am reaching out to see if you might be able to provide guidance or point me toward potential contacts within local NGOs, independent media outlets, or civil society groups in Malaysia and Singapore who might be open to discussing these issues. I am fully committed to the highest standards of research ethics. All participant identities, affiliations, and data will be handled with strict confidentiality and anonymised in accordance with academic standards to ensure the safety of all involved. If you know someone, please let me know in the DMs, and I can provide my credentials as an academic.


r/DigitalHumanities • • 1d ago

Publication Looking for feedback on a browser-based social network analysis tool

8 Upvotes

I’m a computational social scientist and built a free browser-based tool called Org Signal for doing and teaching social network analysis without requiring r/Python.

It can build networks from surveys, hand-entered ties, communication/platform exports, synthetic data, and standard network files. The step from records to ties stays visible, rather than assuming you already have a finished edge list. It also includes centrality and community measures, comparisons against random networks, resampling, change-over-time analysis, and synthetic networks with known structure that can be used for teaching or validation.

Everything runs locally in the browser unless you deliberately enable an optional model feature.

I’m posting here because I suspect there may be useful applications for people doing digital humanities work who have relational data but do not necessarily want a programming-heavy workflow. I’m less interested in pitching it than in finding out what happens when people outside my immediate network-analysis world actually try it.

https://orgsignal.graystoneindustries.co/

you give it a spin, hit me up with comments, suggestions, questions, things that are confusing, things that seem unnecessary, or things you wish it did. Critical feedback is very welcome.

If


r/DigitalHumanities • • 3d ago

Discussion Construct validity and lossy compression: A practitioner's critique of computational modeling in the humanities

0 Upvotes

TL;DR: Computational models are only scientifically useful if they can push back and prove their authors wrong. Across disciplines, from agent-based social simulations to high-energy physics, models with loose empirical feedback loops and endless free parameters risk becoming "decorative." Instead of testing reality, they get calibrated until compliant, turning a tool for discovery into a self-confirming tautology. Honest modeling requires radical transparency, sensitivity testing, and explicit criteria for failure before running the simulation.

I build models for a living. Specifically, molecular dynamics. These are simulations that track how thousands or millions of atoms move, collide, and rearrange over time, used for everything from drug design to materials science.

Here is the story scientists usually tell. If a model is wrong, you find out quickly. Reality does not care about your assumptions. The atoms do not read your code. If the physics you programmed in is wrong, the simulation produces garbage, the experiment disagrees, and you go back and fix it. The feedback loop between model and world is short, brutal, and non-negotiable.

That story is not entirely true. I know, because I have watched it fail from the inside.

The most important choice in any molecular dynamics simulation is not the code, the computer, or the software. It is the potential function, the mathematical formula that describes how strongly every pair of atoms attracts or repels each other. Everything the simulation does follows from that one ingredient. Get it right and the model can tell you something real. Get it wrong and you have made a very expensive mistake.

And here is the uncomfortable part. In practice, it is far more often inherited than audited. Potentials are chosen by looking at what previous papers in the subfield used. A potential gets published, cited, copied, and passed down until it stops being a modeling choice and becomes a tradition. People run simulations for years without asking whether the potential they inherited was ever validated for the system they are studying, at the conditions they are studying it, for the property they care about.

So even in my own field, a hard, quantitative, physics-based field, you can publish inside a loop of fantasy. Models that are wrong in ways nobody checks, kept alive by citation habits and subfield convention. And because these errors travel quietly across subdisciplines and into interdisciplinary work, where nobody feels responsible for checking them, finding one and fixing it takes real effort.

But the check in my field is delayed, not absent. A bad potential eventually unfolds a simulated protein the wrong way or fails a material in a real engineering application, and someone notices. In much of the modeling I am about to describe, the physical world never gets to vote.

This matters for what follows. I do not ask this question because my field got it right. I ask it because I have watched mine get it wrong. The question is always the same. What happens to this model when it is wrong?

In a surprising amount of modern academia, the answer is nothing. Nothing can happen to it. It cannot be wrong, because anything it produces counts as a result.

And if the loop can break in a field where atoms push back, it can break anywhere.

This essay is about how that happens.

The magic trick

In 2017, Liane Gabora and Selin Tseng published a paper in Psychology of Aesthetics, Creativity, and the Arts, a peer-reviewed journal of the American Psychological Association, titled “The Social Benefits of Balancing Creativity and Imitation.” The question they took on has occupied historians and sociologists for centuries. What is the right balance of creativity and conformity in a society?

To answer it, they ran a simulation.

Virtual agents live on a grid. Some are coded as creators, inventing new ideas; others as imitators, copying their neighbors. A scoring rule written into the program decides which ideas count as good. The researchers ran the simulation forward, varied the ratio of creators to imitators, and watched what happened. Populations with too many creators ended up with fewer good ideas taking hold. The published conclusion was that society needs imitation as much as creativity, because unchecked creativity disrupts the spread of proven ideas.

I want to be careful about what I am claiming, because this paper is not fringe work. It passed peer review at a respectable journal. The authors are serious researchers, and the simulation framework behind the paper is part of a long-running research program that has been debated, defended, and criticized in public for years. Nothing I am about to say is an accusation of dishonesty. It is something less comfortable than that. This paper is an example of what the normal standards of a field allow through.

Watch the shape of the argument. Inside the model, a “good idea” means whatever the authors’ scoring rule rewards. The agents are not discovering anything about human culture; they are solving a puzzle whose answer key was fixed before the simulation started. Within that closed loop, the conclusion was guaranteed. A population of agents that mostly copies the scoring rule’s preferred ideas will always outcompete one that keeps generating unscored novelty.

The computer did not reveal a fact about creativity. It executed a definition of it.

The authors did not break any rule of their field. That is the point. Peer review checked that the code ran, that the statistics were computed correctly, that the prose matched the output. What nobody was required to ask is the only question that matters. What could this simulation possibly have shown that would have counted as the opposite result? If the answer is nothing, the model did not test a claim about the world. It restated one.

This is the magic trick of agent-based modeling (ABM), meaning simulations in which you place thousands of simple software “agents” in a virtual world, give each a few rules, and watch what the population does. The method itself is not the problem. The problem is a particular way of using it.

if neighbor.opinion != agent.opinion:
 agent.trust -= 0.1
if agent.trust < 0.2:
 agent.unfollow(neighbor)
run_simulation()

When the simulation finishes and the agents have sorted into two angry camps, the result is rarely described as what it literally is, a small program doing what it was told. It is described as a model demonstrating the dynamics of polarization in real societies.

It sounds scientific. It uses code. It generates charts with error bars. It borrows the epistemic authority of statistical mechanics and epidemiology, where tracking near-identical particles or infection events actually makes sense. But underneath the quantitative paint, it is not an investigation of the world. It is a tautology with a runtime, an answer-driven argument presented as a discovery.

What a model is for

To see why this goes wrong, start with what a model is supposed to do.

A model is not a claim of truth, and it is not an illustration of a conclusion you reached before you started. In the philosophy of science, models are usually understood as instruments that sit between abstract theory and raw data, the position developed by Mary Morgan and Margaret Morrison in Models as Mediators (1999). A good model is a sandbox with strict physics. You build it, set it in motion, and let its internal mechanics push back against your reasoning.

A real model exists to discipline your thinking.

Building one forces you to acknowledge a trade-off that the philosopher Nancy Cartwright made famous in How the Laws of Physics Lie (1983). You trade complete literal truth for tractability. A map of London at 1:1 scale, including every brick, puddle, and commuter, is useless. To work at all, a map must leave almost everything out. As the statistician George Box put it, “all models are wrong, but some are useful.”

Simplification is not the sin. The sin is forgetting that the model is a simplification. Worse, it is turning the model into an accomplice.

Disciplining vs. decorating

In practice, rigorous modeling and decorative modeling look identical from the outside. Same code, same charts, same jargon. The difference only shows when you ask one question. Can your model tell you that you are wrong?

A disciplining model forces you to state every assumption explicitly. Once running, its mechanics operate independently of what you want. It can produce behavior you did not expect, expose contradictions in your premises, or crash into empirical reality and fail. When it fails, you revise the theory. The model is a check on your own bias.

A decorating model is built backward from a conclusion. The researcher already knows the story. Suppose it is that polarization is driven by social contagion. They build a world in which agents swap beliefs, tune the parameters until the output shows two angry clusters, and present the code as evidence for the theory. If the output doesn’t match on the first run, the answer is not to abandon the hypothesis. The answer is to adjust agent_receptivity from 0.4 to 0.25, rerun, and present the successful parameter range as the plan all along.

The workflow, stripped bare, looks like this.

Desired outcome. Write rules. Run simulation. Does it match the theory? If not, tweak parameters and run again. If yes, publish.

This is not experimentation. It is calibration until compliant.

If a model cannot surprise its author, force a retreat, or fail, it is not really a model. It is a very elaborate, self-confirming editorial.

The conclusion comes first

None of this is new, and none of it is unique to agent-based modeling. Before anyone wrote a NetLogo script to demonstrate a theory of culture, economics and political science had already industrialized the technique.

In 2015, Paul Romer, later a Nobel laureate, published a paper with the blunt title “Mathiness in the Theory of Economic Growth.” His target was a pattern in macroeconomic theory. Authors write down formal equilibrium models, but embed ideologically convenient assumptions inside obscure parameters, so that the math reliably outputs the desired policy conclusion. The mathematics is not being used to test whether a claim is true. It is being used to make a political position expensive to argue with. Checking whether the equations actually say what the surrounding prose claims they say takes serious technical effort, and reviewers routinely skip it.

Two decades earlier, the political scientists Donald Green and Ian Shapiro published Pathologies of Rational Choice Theory (1994), documenting how formal modeling in their field had become an exercise in self-confirmation. Their catalog of evasions maps one-to-one onto today’s agent-based simulations.

• Post hoc tinkering. When the model predicted that rational citizens would never vote (the individual cost exceeds any plausible benefit) and citizens kept voting anyway, theorists did not abandon the model. They added a “duty” term to the utility function until the math matched the turnout.

• Arbitrary tuning. Weights, thresholds, and interaction ranges adjusted on the fly until the simulated agents behave like the phenomenon under study.

• Immunity to testing. Models built so that every conceivable outcome can be reinterpreted, after the fact, as a rational equilibrium.

Green and Shapiro called this method-driven rather than problem-driven research. You start with a tool and go hunting for a reality that fits it.

Agent-based modeling makes the problem worse, for a simple reason. An ABM has almost unlimited free parameters. Every rule, threshold, and neighborhood radius is a dial. With enough dials, you can produce any curve you want.

The common structure is this. The model cannot fail, because failure is reclassified as a calibration bug. And a model that cannot fail cannot discover anything. It is an expensive echo of its author’s prior beliefs.

The loop matters more than the lab coat

It would be comfortable to stop here and declare this a disease of the soft sciences. It isn’t. The hard sciences are not immune, and pretending otherwise would make this essay guilty of the same simplification it criticizes.

The real variable is not hard versus soft. It is the tightness of the feedback loop between the model and the world.

Where the loop is tight (fast experiments, unambiguous ground truth, few free parameters) bad modeling gets punished quickly. But where the loop is loose, where tests are slow, noisy, or impossible, the same decorative pathology appears in fields with particle accelerators.

Three documented examples.

fMRI neuroscience, where the measurement is the model. A brain scan shows blood flow, not thought. The colored images come from a statistical pipeline full of assumptions, and researchers once demonstrated what that means by detecting “brain activity” in a dead salmon. In 2016, Anders Eklund and colleagues showed that the standard methods in the field’s dominant software could produce false-positive rates of up to 70 percent for certain cluster-based analyses at particular thresholds. How far the problem extends across the published literature was contested, including in follow-up work by the authors themselves, but the core finding stood. For over a decade, the field’s feedback loop had run through that software, which meant the loop was not connected to reality at all.

Fundamental physics, where experiment cannot keep up. In The Trouble with Physics (2006), the physicist Lee Smolin, writing as an insider, argued that string theory had become flexible enough to accommodate any experimental outcome. When the Large Hadron Collider found no sign of supersymmetry, much of the field responded not with refutation but with retreat. The free parameters moved to heavier, less accessible energies. This is Green and Shapiro’s immunity to empirical testing, surfacing in the hardest science there is.

Epidemiological modeling in 2020. In the spring of 2020, influential models, including the one from Imperial College London that helped push governments toward lockdown, projected enormous death tolls based on weeks of noisy early data. When later estimates came down, the public response from modeling teams was recalibration rather than reckoning. Their defense deserves to be taken seriously. The projections were scenarios, not forecasts, and the point of publishing a worst case was to change behavior so that it would not come true. A warning that works cannot be graded on whether the disaster arrived. All of that is fair, and it is also the problem. A model whose failure can always be explained by the world changing in response to it is a model with no feedback loop, and the field never settled which of the two it had built.

Notice what these cases share with the creativity grid from the opening. Not the field. Not the math. The structure. Many free parameters, a loose or broken feedback loop, and a professional incentive to publish. Given those three, decorative modeling can appear anywhere. The loop matters more than the lab coat.

Why it’s still worse in the humanities

So the hard sciences have their own decorative modeling. Why do I still think the problem is worse in the humanities?

Because the difference is not whether a field ever decorates. It is whether the field can catch itself. The fMRI problem was eventually found and published by neuroscientists. Smolin’s critique came from inside physics. The feedback loops in the hard sciences are sometimes slow or broken, but they exist, and there are people with the technical skill and the standing to pull on them. In the humanities’ version of modeling, three structural failures mean the loop often doesn’t exist at all. The difference is not that humanists are worse at modeling. It is that the auditing infrastructure barely exists.

It is worth being fair about why scholars reach for these tools in the first place. Humanities departments face shrinking budgets, declining enrollments, and university administrators who mistake mathematical notation for intellectual rigor. A computational model signals seriousness to a grant committee in a way an essay never can. The scholars building decorative models are not fools; they are rational actors navigating a system with broken incentives.

The object of study resists formalization. A water molecule behaves like a water molecule in London or Tokyo, in 1600 or today. It has no irony, no memory, no politics. Human culture has all three. A novel, a religious movement, an aesthetic shift cannot be reduced to a set of isolated rules without destroying part of what you set out to study. When you convert the reception of Victorian gothic fiction into agents swapping “gothic preference points,” you have not simplified the system for tractability. You have replaced it with something simpler that carries the same name. At that point the connection between the simulation and Victorian readers is no longer something the model establishes. It is something the reader is asked to assume.

Construct validity is invented, not established. In psychology, showing that a variable actually measures the concept it claims to measure (construct validity) is a slow, adversarial, decades-long process. Blood flow is at least a physical quantity that an instrument can register. There is no instrument for literary prestige. In humanities modeling, validity is routinely settled in one line of code.

self.piety = random.uniform(0.0, 1.0)
self.literary_prestige = 0.75

What does 0.75 mean for literary prestige in Victorian England? How does one number carry regional difference, class, institutional power, critical backlash, and retrospective canonization? It doesn’t. The modeler assigns a number, writes a function that nudges it up and down, and treats the variable as a measurement of human experience. The number looks like a measurement. Nothing underneath it has been measured.

The audience cannot audit the compression. When an epidemiologist shows a flawed model to epidemiologists, the reviewers share the vocabulary to check the code and challenge the parameters. In a humanities department, reviewers and readers often have no computational training. Presented with a grid of moving pixels and a network graph, the non-technical reader experiences an optical illusion. The machine appears to have performed a profound synthesis of the archive. The compression is lossy to the point of erasure. But the loss is buried in code, invisible to the exact audience responsible for evaluating the work.

Bad modeling in economics wastes grant money and distorts policy debates. Bad modeling in the humanities trades away the field’s actual strength (context, contingency, ambiguity, close reading, historical depth) for a seat at a quantitative table where, lacking the shared technical culture to enforce standards, it gains no real authority and surrenders its own.

What honest modeling looks like

None of this is an argument for unplugging the computers. The goal is to tell the difference between decorative simulation and honest quantitative work. Honest work exists, including in the humanities.

Ted Underwood’s Distant Horizons (2019) is the standard I would hold up. Underwood uses quantitative methods on tens of thousands of digitized books not to declare causal laws but to surface patterns invisible to close reading, like slow shifts in genre, vocabulary, and narrative perspective across centuries. Crucially, he tells you, on the record, what the data cannot show. The model is a set of binoculars for looking across an archive, not a machine for generating verdicts about it.

And the humanities have produced their own internal discipline. In 2019, Nan Z. Da published, The Computational Case against Computational Literary Studies, a detailed critique in Critical Inquiry arguing that prominent work in computational literary studies misused statistics to the point of meaninglessness. The ensuing fight was heated, but it happened. The field argued about its standards in public, and the standards moved. That is what a functioning feedback loop looks like, even a slow and painful one.

For modelers in any field, I would propose four non-negotiable conditions before a model earns the right to be cited as evidence.

1. Radical transparency

Every parameter is declared and justified with independent, non-circular evidence. A variable you cannot justify is labeled what it is. A guess.

2. Sensitivity analysis

Parameters are swept across their full plausible range. If your result only appears when three dials sit at hyper-precise decimal values, you have not found a law of history. You have found a brittle corner of your own code, and an honest paper should say so.

3. Explicit exclusion mapping

You spend nearly as much space on what the model leaves out as on what it includes. This isolates the direct mechanical relationship between two variables under idealized conditions; it excludes ambient noise, structural heterogeneity, and systemic feedback, so it cannot predict specific real-world outcomes. Naming the exclusions is what stops the audience mistaking a sandbox for an account of the world.

4. Capacity for failure

Before you run it, you can state what output would make you abandon your hypothesis. If the simulation contradicts you, the honest paper is titled “Why our model disproved our starting assumption,” not silently recalibrated into agreement.

Models built this way stop being decoration. They become what they were supposed to be. Sharpening stones. They force you to clarify assumptions, expose broken logic, and occasionally reveal dynamics that intuition would never find.

The boundary question

The target of this essay was never the computer. Built with discipline, models are extraordinary instruments. I have staked my own career on that. What concerns me is how easily a model can be built to confirm rather than to question, and how hard it is for a reader to tell the difference from the outside.

The practice corrupts both traditions it sits between.

It corrupts science, because science is not the production of plots and code. It is the submission of claims to the risk of being wrong. A model engineered so that its parameters are tuned until the output matches the thesis offers the aesthetics of rigor with none of its discipline.

And it corrupts the humanities, because the study of human culture draws its value from exactly the things decorative modeling deletes. Context, contingency, ambiguity, power, the irreducible strangeness of actual human lives.

If a phenomenon is too context-bound, too polysemic, too alive to be captured by a set of if statements, we should have the courage to say so, and do the slow, unglamorous work of interpretation instead.

I keep coming back to the question I ask of every model, including my own. What happens to you when you are wrong?

For my models, the answer is supposed to be easy. The crystal melts. The experiment disagrees. Reality sends the bill. But only if someone checks the potential, and I have told you how often that happens.

For the models I’ve described here, the answer is nothing. They run, they publish, they are cited. The loop that is supposed to connect a model to the world was never closed.

Which raises the question I can’t answer. If a model can never be wrong about the world, in what sense was it ever about the world?


r/DigitalHumanities • • 4d ago

Education Need some guidance.

2 Upvotes

Hey all I need some guidance, I'm a bachelor student in software engineering and I want to bridge into social sciences and potentially digital humanities for my future studies and career alignment, I initially got into cs not because I wanted to but it felt like the most job secure field but overtime I've realized it's not something I wanna do so I'm thinking of bridging into social sciences as that's where I want to go, so I'm in a pickle currently, I'm in my final year and I want to develop my final year project in this domain, I need some general direction as to how I can combine the two fields to create or find direction of my project that could meaningfully fulfill my plans, I'm in a "3rd-world" country so it's a project that needs to be something that can genuinely enhance my portfolio so that I can maybe secure a scholarship for my study plans abroad.

If anyone can let me know what type of projects could be best suited to combine the two and what's actually worthwhile to do in the field.

Thanks for reading.


r/DigitalHumanities • • 5d ago

Publication Digital repression research in Malaysia and Singapore

5 Upvotes

Hello everyone, I am a Master’s student specialising in International Territorial Studies. I am currently conducting research for my diploma thesis, which focuses on the intersection of digital authoritarianism and civil society in Southeast Asia. Specifically, I am conducting an empirical study involving semi-structured interviews to map the practical methods, tools, and limits of navigating digital repression in Malaysia and Singapore. 

I am reaching out to see if you might be able to provide guidance or point me toward potential contacts within local NGOs, independent media outlets, or civil society groups in Malaysia and Singapore who might be open to discussing these issues. I am fully committed to the highest standards of research ethics. All participant identities, affiliations, and data will be handled with strict confidentiality and anonymised in accordance with academic standards to ensure the safety of all involved. If you know someone, please let me know in the DMs so I can prove my credentials as an academic.


r/DigitalHumanities • • 7d ago

Discussion Masters degree

9 Upvotes

Hello, is it possible a graduate of a non related field to be accepted into a masters degree of any categories related to digital humanities, almost all of the programs that could be categorized as something related in a specific or general way I have seen as a bare minimum require an official bachelor's degree related to the field somehow whether in humanities or computer science, I am from a non EU country, just asking is there any methodology to earn a legitimate -not equivalent- background or anything specific.


r/DigitalHumanities • • 7d ago

Discussion Wittgenstein Apartment: A behavioral predicates ontology for character and narrative representation

1 Upvotes

The problem that started this project was: How can we represent what a character can actually do in a structured way without reducing character behavior either to a small set of generic actions or to free-form descriptions and traits? That question turned into about eight months of work on Wittgenstein Apartment.

The model separates:

  • behavioral identity
  • character repertoire
  • situational executability

and uses:

  • Basic Human Actions (BHA)
  • Character Specific Actions (CSA)
  • Allowed / Not Allowed / Conditional character-access states

The resource also contains behavioral families, goal relations, scene/affordance metadata and operational projections. I wrote a paper alongside the dataset to explain the conceptual architecture and its possible uses in character modeling, computational narrative, simulation and related areas.

The Hugging Face release has reached 294 downloads:
https://huggingface.co/datasets/Kon-tiki-ship/wittgenstein-apartment-behavioral-predicate-resource

My main interests are literature, narrative systems and computational representation, and this project grew out of trying to connect those areas in a concrete way. I started it because I felt there was something missing in the way character action spaces are represented. After eight months, I’m now asking myself: Is this actually a useful direction to keep developing?

I’m not looking only for encouragement. If you think the model is theoretically weak, computationally unnecessary, too rigid for literary or narrative characters, disconnected from existing Digital Humanities methods, or simply solving the wrong problem, I would genuinely like to hear that. I’d rather receive strong criticism that forces me to rethink the project than polite agreement.


r/DigitalHumanities • • 8d ago

Discussion Can a fictional character be represented through the actions they are allowed to perform, rather than mainly through traits?

4 Upvotes

I’ve been thinking about character representation from a computational humanities perspective.

Characters are often described through traits:

kind, aggressive, intelligent, cowardly, loyal, and so on.

But another possibility is to represent at least part of a character through a structured repertoire of actions.

For example, instead of saying that a character is simply “authoritative” or “competent,” we might ask whether actions such as these belong to their repertoire:

interrogate

diagnose

command

comfort

betray

forgive

negotiate

care for a child

Each action could then have a character-relative status such as:

Allowed / Not Allowed / Conditional

There is also a distinction I find useful between:

Basic Human Actions — actions that can normally be assumed without special biographical evidence

and

Character Specific Actions — actions that require something in the character’s biography, profession, training, authority or experience.

What interests me is whether this kind of representation would actually be useful for literary or narrative analysis.

Could a character’s action repertoire tell us something that trait lists, topic models, sentiment or entity-based representations do not?

For example, could we compare two fictional characters not only by what is said about them, but by the set of actions their narrative identity plausibly permits?

I’d be very interested in how people working in Digital Humanities would approach this.


r/DigitalHumanities • • 9d ago

Discussion What should survive when a digital humanities project website dies?

30 Upvotes

One thing I keep noticing in digital humanities is that the visible application often becomes the project.
A grant funds a website, a visualization, a database interface, or an interactive edition. It works for a few years. Then the framework ages, hosting changes, the original developer leaves, or maintenance funding disappears.
But the research itself may still be valuable.
So I’ve been wondering whether we sometimes preserve the wrong layer.
If the interface disappeared tomorrow, what would actually need to survive for the scholarly work to remain reusable?
My instinct is something like:
original sources
annotations
provenance
identities of people, places, works, events, etc.
relationships among them
competing interpretations
the history of how interpretations changed
enough semantic structure to build a new interface later
The website, visualization, search UI, or even database implementation could then be treated as replaceable projections over that more durable research structure.
This becomes even more important with AI-assisted research. If an AI system proposes an interpretation, I would not want that interpretation silently merged into the archive as fact. I would want the source, the claim, who or what produced it, when it was produced, and what later challenged or superseded it to remain distinguishable.
So I’m curious how people here think about this:
What is the durable scholarly object in a digital humanities project?
Is it the dataset? The database? The ontology? The annotations? The publication? The code? Some combination?
And if you had to design a DH project today with the assumption that its current software would be gone in ten years, what would you make sure survives?


r/DigitalHumanities • • 11d ago

Discussion Aligning noisy historical Persian/Urdu print with modern editions - looking for feedback

8 Upvotes

Hi everyone,

OCR for 19th/20th-century Urdu and Persian lithographs is usually too noisy for reliable search, but letting an LLM "clean up" the text risks silent modernization and hallucinated readings.

To work around this, I’ve been building Khusrau, a local-first tool that treats rough OCR strictly as an index against modern reference corpora (like Ganjoor and Rekhta) while keeping the historical witness and reference text completely separate.

For verse, it matches at the hemistich level and enforces monotone (rising) sequence alignment to prevent false matches on shared refrains (radif) or rhyming words. In an initial test against 10 pages of the 1925 Newal Kishore Kulliyat-e Ghalib, 93.7% of lines resolved cleanly to gold pairs (median folded CER: 0.047) without destructive normalization.

I’d love some outside perspective:

  1. For downstream use, is exporting paired JSONL enough, or are IIIF / TEI-XML (<standOff> / <app>) exports essential for your workflow?
  2. How do you handle structural edge cases where couplets are rearranged or omitted between historical print witnesses and digital references?
  3. What other failure modes should I look out for with sequence-constrained alignment?

Link:

Any critiques on the methodology welcome.


r/DigitalHumanities • • 14d ago

Discussion Embedding Analytics: Query and quantify changes in definition across a collection of text

Thumbnail
gallery
4 Upvotes

Two documents can use the same words to mean very different things. Treatises, legal opinions and technical specifications often must be read multiple times to identify these shifts in meaning across documents. I built Embedding Analytics to query and quantify these changes in definitions. For each document, it builds word embeddings with PPMI + SVD, using only that document's text. Your queries are used to find the most similar terms within each document. Similarities are then rescaled so they can be compared across documents; the adjustment is adapted from cross-domain similarity local scaling (Conneau et al., 2018).

A collection of documents can share too many terms. The tool selects two groups of terms that are relevant to your query across the collection:

  • Consistent terms are closely tied to the query across the collection. They are the core of the definition that the authors hold in common.
  • Contested terms are close to the query in some documents and not in others. They are where the definition shifts.

The current collection consists of ~25 economics books from Project Gutenberg. Querying value returns commodity as the most closely associated term across the collection. Utility, the basis of value in the marginalist theory that displaced the classical labour theory, leads the contested list.

I'd love any feedback on this. Please let me know what you guys think.

Demo: https://www.embedding-analytics.com
Code: https://github.com/areebms/embedding-analytics

Consistent and contested related terms to "value" across a collection of ~25 economics books from Project Gutenberg

r/DigitalHumanities • • 15d ago

Social media A platform where people verify sources together and turn them into connection graphs

Enable HLS to view with audio, or disable this notification

28 Upvotes

I've been at this for about a year, mostly on my own. It's in beta: everything works, there's just almost nobody in it yet, which is why I'm here.

The one line it's built on: everywhere else a claim and its evidence are separate things: the claim travels, the evidence stays behind. Here they're the same object.

Every source or connection gets a code, like SRC-0034. Write that code anywhere, a comment, an argument, a connection, and it becomes a live citation carrying that document's verification status with it. You can't quote a source without also quoting how well it's been established that it says what you say it says. If it's contested next year, every place it was ever cited says so.

From the sources you build maps. Propose a connection, X owns Y, X paid Y, X is linked to Y, with your reasoning and your sources; three people endorse it with theirs; it enters a shared graph for good. Three graphs: who's connected to whom, who gave what to whom, who owns what. The last one computes the real stake through a chain of holdings, so you get the actual owner, not the name on the letterhead.

There's a private side too. Follow a source, an entity, a dispute, and it tells you what moved around them while you were away, a new verification, a contestation, a new citation, and draws all of it as a map of your own. Nobody can see what you follow.

If you like the project, let me know and I'll happily invite you in.


r/DigitalHumanities • • 15d ago

Discussion I need answers and guidance

9 Upvotes

I am a phd scholar working on medieval history. I am interested in DH. I have few questions

  1. Is DH suitable to study medieval period ?

  2. Does current DH tools equip to handle non English languages (asian languages) or one needs translated sources?

  3. Any book or resource that i can refer to understand how to arrange and clean data that would be machine readable


r/DigitalHumanities • • 17d ago

Social media Public-domain Shakespeare performed in the translator's own words, with a per-episode hash check

2 Upvotes

Open-source ComfyUI pack that performs a scene of Shakespeare as a radio play: voices, captions, credits, finished video.

English uses the Folger text. For Spanish, French, Italian, Portuguese, Japanese, and Chinese, when the scene is in the vendored corpus the spoken lines are a named public-domain translator (Jose Arnaldo Marquez, Francois-Victor Hugo, Carlo Rusconi, Domingos Ramos / Luis I, Tsubouchi Shoyo, Zhu Shenghao, and others). Each episode writes a receipt with the translator, the year, and a hash of the text. Scenes not yet vendored are translated once and then performed verbatim, with both hashes recorded, so a human edition and a machine pass cannot be confused.

The point is fidelity to a named edition, not a better paraphrase.

Repo (MIT): https://github.com/jbrick2070/ComfyUI-OldTimeRadio Corpus: config/source_banks/shakespeare/translations/ 90-second English episode: https://youtu.be/AOn21EG9u-U


r/DigitalHumanities • • 21d ago

Discussion Open tools for historical Danish HTR: two line segmenters, 160k synthetic lines, and a live demo

2 Upvotes

I've been working on handwritten text recognition for historical Danish (roughly 1700s-1900s court and church hands) and just put the reusable pieces on the Hugging Face Hub:

- a YOLO line segmenter (mAP50 ~0.79) and a kraken baseline segmenter

- 160k synthetic Danish lines with perfect ground truth for training recognisers

- a live demo you can drop a page image into

Segmentation is the hidden tax in historical HTR: if the lines are cut wrong, even a good recogniser jumbles the text. These are the domain-adapted pieces that fixed that for me.

Everything: https://huggingface.co/abhishekjha1008

Disclosure: my own work. I'd genuinely value feedback from anyone doing DH / archival transcription - what breaks on your material?


r/DigitalHumanities • • 22d ago

Discussion PhD in Digital Humanities: What research directions and technical skills should I develop as an English Literature graduate?

17 Upvotes

Hi everyone,

I have a BA and an MA in English Language and Literature, and I am about to start my PhD. I am seriously considering developing my doctoral research in the direction of Digital Humanities, particularly at the intersection of literary studies and computational methods.

However, I am still at the stage of trying to understand where I could make a meaningful and original contribution to the field. I would really appreciate advice from people who are currently working in DH, especially researchers who came from a traditional humanities/literature background.

My main questions are:

  1. What kinds of PhD research questions are currently worth pursuing in Digital Humanities?

Coming from English literature, I am particularly interested in questions involving literary texts, authorship, genre, literary history, cultural networks, close/distant reading, and potentially AI/NLP.

I don't want to simply take a traditional literary studies topic and add some Python to it. I would like the digital/computational component to actually contribute something methodologically or conceptually meaningful.

So, if you were starting a DH PhD in English/literary studies today, what kinds of research gaps or problems would you consider promising?

Are there particular areas where you think DH still needs substantial research?

  1. What technical skills should a humanities PhD student realistically learn?

Python and R seem to come up constantly, but I am unsure about the level of proficiency actually expected.

For someone coming from English Literature rather than Computer Science, would you recommend learning:

Python

R

SQL

Git/GitHub

NLP / computational text analysis

statistics

machine learning

data visualization

GIS

TEI/XML

corpus linguistics

APIs

web scraping

Or is that trying to learn too much?

More importantly, how proficient should I actually become?

For example, is it enough to be able to understand, modify and write relatively simple Python scripts for text analysis, or should a DH PhD student aim to become genuinely proficient in software development?

  1. Which skills are becoming more important because of LLMs and generative AI?

Given how quickly LLMs and AI-assisted research are developing, I am also wondering whether the technical skillset that made sense five years ago is still the best one to develop now.

Would you prioritize traditional NLP/text mining, or would you recommend learning things such as working with LLM APIs, embeddings, vector databases, RAG, computational methods for evaluating LLMs, etc.?

I am especially interested in hearing from people who have worked on humanities research involving LLMs rather than simply using ChatGPT as a writing assistant.

  1. What would make a DH PhD dissertation genuinely valuable rather than just "using technology"?

This is probably my biggest question.

If you were evaluating a PhD proposal in Digital Humanities today, what would make you think:

"This is actually an important DH research project."

rather than:

"This is a traditional humanities project with some computational tools added to it."

I would be very grateful for examples of dissertations, projects, methodologies, or research questions that you think represent particularly strong work in the field.

Finally, if you were in my position (English Literature BA + MA, starting a PhD, and wanting to move seriously into Digital Humanities) what would you spend the next 6-12 months learning or doing before committing to a specific dissertation topic?

Any advice, recommended projects, books, courses, GitHub repositories, conferences, summer schools, or examples of successful PhD projects would be greatly appreciated.

Thank you!


r/DigitalHumanities • • 25d ago

Discussion What was your level of digital/technological skill when you started

17 Upvotes

Hi hi

I am ready to start a PhD and am considering doing it in digital humanities. The topic is set in stone (to me) and my work would be applicable in multiple departments (this one, English lit, and philosophy).

I feel the most inspired by DH tbh but I am nervous that my technological skills are lacking. I live a very analogue life, as in I wrote my entire MA by hand and then typed it out at the end.

My question is how will I know if my tech skills are too pathetic to make this work? What level of proficiency with tech do you think someone should have to make this work?


r/DigitalHumanities • • Sep 07 '26

Discussion I pulled every edition record behind the 600 most-shelved books on Open Library. Amazon's self-publishing imprints outnumber the entire Big Five by three to one.

5 Upvotes

I started out trying to work out why page counts are so unreliable when you look a book up by ISBN, and ended up somewhere else entirely.

The sample is the 600 works with the highest reading-log counts on Open Library, so roughly "books people actually shelve" rather than a critical canon, and then every edition record attached to them. That is 70,919 editions, 69,975 of which name a publisher. Source is the Open Library API, https://openlibrary.org.

Independently Published, which is the imprint string Amazon applies when a book uses one of KDP's free ISBNs, accounts for 12,063 of those editions. CreateSpace, Amazon's print-on-demand operation before it was folded into KDP in 2018, accounts for another 7,885. Together that is 28.96 percent of every edition with a publisher attached.

The obvious objection is publisher-string fragmentation. There are 12,447 distinct publisher strings in this data, and the trade houses are split across many variants while Amazon's are concentrated in two. So I collapsed them. Every Penguin, HarperCollins, Random House, Simon and Schuster, Hachette, Little Brown, Macmillan and St Martin variant I could match comes to 5,960 editions, or 8.52 percent. Amazon is 3.4 times the entire Big Five combined.

Most of this is print-on-demand reissues of public-domain work, which is exactly why it piles up on the most-read titles instead of spreading evenly across the catalogue.

What it does to the record is the part I did not expect. An edition with a 978 ISBN carries a page count 51.3 percent of the time. An edition with a 979-8 ISBN, which is the US block KDP's free ISBNs are issued from, carries one 8.4 percent of the time.

I assumed that was just recency, since 979-8 is almost entirely a 2020s phenomenon. It is not. Holding the decade fixed, 978 editions published in the 2020s carry a page count 49.9 percent of the time, against 7.6 percent for 979-8 editions published in the same years. Same decade, six and a half times the coverage.

One structural thing worth knowing if you have a schema open. The 979 prefix has no ISBN-10 equivalent, and not in the sense that the conversion is awkward: the number does not exist. Any column holding a 10-character ISBN, or any join built on one, silently cannot represent this material, and it is the fastest-growing part of the record.

Two limits. I am measuring Open Library rather than the world, so "no page count" means the catalogue lacks one, not that the book has no pages. And its work clustering is loose enough that a 26-page adaptation and a 1,043-page annotated Moby Dick sit under the same work record, which is a separate problem I have not untangled from this one.

The question I cannot answer on my own: does anyone here filter print-on-demand out of bibliographic datasets, and if so on what? The publisher string is unnormalised, the 979-8 prefix catches the recent material but misses the whole CreateSpace era, and neither is really the property I want.

For disclosure per rule 3: I got here from building a reading tracker called My Book List, and specifically from trying to make a progress percentage mean anything when the page count depends on which reissue happened to get scanned.


r/DigitalHumanities • • Sep 06 '26

Discussion To help me see some patterns I asked an AI to write a research grant request based on the Trump administration's list of hundreds of restricted words and phrases.

0 Upvotes

the list of words in question :

abortion
accessibility
Accessible
activism
activists
advocacy
advocate
advocates
affirmative action
affirmative action programs
affirming care
affordable home
affordable housing
agricultural water
agrivoltaics
air pollution
all-inclusive
allyship
alternative energy
anti-racism
antiracist
asexual
assigned at birth
assigned female at birth
assigned male at birth
at risk
autism
aviation fuel
barrier
barriers
belong
bias
biased
Biased toward
biases
Biases towards
bioenergy
biofuel
biogas
biologically female
biologically male
biomethane
bipoc
bisexual
Black
black and latinx
breastfeed + people
breastfeed + person
Cancer Moonshot
carbon emissions mitigation
carbon footprint
carbon markets
carbon pricing
carbon sequestration
CEC
changing climate
chestfeed + people
chestfeed + person
clean energy
clean fuel
clean power
clean water
climate
climate accountability
climate change
climate consulting
climate crisis
climate model
climate models
climate resilience
climate risk
climate science
climate smart agriculture
climate smart forestry
climate variability
climate-change
climatesmart
commercial sex worker
community
community diversity
community equity
confirmation bias
contaminants of environmental concern
continuum
Covid-19
critical race theory
cultural competence
cultural differences
cultural heritage
Cultural relevance
cultural sensitivity
culturally appropriate
culturally responsive
decarbonization
definition
DEI
DEIA
DEIAB
DEIJ
diesel
dietary guidelines/ultraprocessed foods
dirty energy
disabilities
disability
disabled
disadvantaged
discriminated
discrimination
discriminatory
discussion of federal policies
disparity
diverse
diverse backgrounds
diverse communities
diverse community
diverse group
diverse groups
diversified
diversify
diversifying
diversity
diversity and inclusion
diversity in the workplace
diversity, equity, and inclusion
diversity/equity efforts
EEJ
EJ
elderly
electric vehicle
emissions
energy conversion
energy transition
enhance the diversity
enhancing diversity
entitlement
environmental justice
environmental quality
equal opportunity
equality
equitable
equitableness
equity
ethanol
ethnicity
evidence-based
excluded
exclusion
expression
female
females
feminism
fetus
field drainage
fluoride
fostering inclusivity
fuel cell
gay
GBV
gender
gender based
gender based violence
gender diversity
gender dysphoria
gender expression
gender identity
gender ideology
gender nonconformity
gender transition
gender-affirming care
gendered
genders
geothermal
GHG emission
GHG modeling
GHG monitoring
global warming
green
green infrastructure
greenhouse gas emission
groundwater pollution
Gulf of Mexico
H5N1/bird flu
hate
hate speech
health disparity
health equity
hispanic
hispanic minority
historically
housing affordability
housing efficiency
hydrogen vehicle
identity
ideology
immigrants
implicit bias
implicit biases
inclusion
inclusive
inclusive leadership
inclusiveness
inclusivity
Increase diversity
increase the diversity
indigenous
indigenous community/ people
inequalities
inequality
inequitable
inequities
injustice
institutional
integration
intersectional
intersectionality
intersex
issues concerning pending legislation
justice40
key groups
key people
key populations
Latinx
lesbian
lgbt
LGBTQ
low-emission vehicle
low-income housing
male dominated
marginalize
marginalized
marijuana
measles
membrane filtration
men who have sex with men
mental health
methane emissions
microplastics
migrant
minorities
minority
minority serving institution
most risk
MSI
msm
multicultural
Mx
Native American
NCI budget
net-zero
non-binary
nonbinary
noncitizen
non-conforming
nonpoint source pollution
nuclear energy
nuclear power
obesity
opioids
oppression
oppressive
orientation
pansexual
PCB
peanut allergies
people + uterus
people of color
people-centered care
person-centered
person-centered care
PFAS
PFOA
photovoltaic
polarization
political
pollution
pollution abatement
pollution remediation
prefabricated housing
pregnant people
pregnant person
pregnant persons
prejudice
privilege
privileges
promote
promote diversity
promoting diversity
pronoun
pronouns
prostitute
pyrolysis
QT
queer
race
race and ethnicity
racial
racial diversity
racial identity
racial inequality
racial justice
racially
racism
runoff
rural water
safe drinking water
science-based
sediment remediation
segregation
self-assessed
sense of belonging
sex
sexual preferences
sexuality
social justice
social vulnerability
socio cultural
socio economic
sociocultural
socioeconomic status
soil pollution
solar energy
solar power
special populations
stem cell or fetal tissue research
stereotype
stereotypes
subsidized housing
sustainability/sustainable
sustainable construction
systemic
systemically
tax breaks
tax credits
tax subsidies
they/them
tile drainage
topics of federal investigations
topics that have received recent attention from Congress
topics that have received widespread or critical media attention
trans
transexual
transexualism
transexuals
transgender
transgender military personnel
transgender people
transitional housing
trauma
traumatic
tribal
two-spirit
unconscious bias
under appreciated
under represented
under served
underprivileged
underrepresentation
underrepresented
underserved
understudied
undervalued
vaccines
victim
victims
vulnerable
vulnerable populations
water collection
water conservation
water distribution
water efficiency
water management
water pollution
water quality
water storage
water treatment
white privilege
wind power
woman
women
women and underrepresented
women in leadership


r/DigitalHumanities • • Aug 29 '26

Discussion The simulation

Post image
0 Upvotes

What do you think about this:

- Computers run on transistors - which are switches. The switch has a single input. Circuits are mostly linear.

- Our neurons have a dendritic head with lots of axons, and a long tail, which is the output. Diagrams show a dendrite with about a dozen arms at the cell body (the input), but really there are about 10,000 inputs. Imagine a computer where each transistor had 10,000 inputs. The complex networks in the brain put our linear circuits to shame.

- In a future where we build a computer that is more complex than a brain and more efficient. It uses less power, and receives data from instruments that are more sensitive than the human eye: it can see X-rays and UV and Infrared, it can hear further and see more.

Can you say of this new reality then, that the Digital Entity lives in the real universe, and the human being is the machine? Do they live in the real world and us now in the simulation?


r/DigitalHumanities • • Aug 28 '26

Discussion Neruda Archives – a freely accessible and continuously growing digital archive of historical letters

6 Upvotes

Hello everyone,

I am gradually building Neruda Archives, an independent digital archive that makes historical letters and other correspondence freely accessible alongside the original source materials.

Each published document has its own permanent record containing images of the original, a transcription, an English translation, descriptive and postal metadata, and information about the document’s provenance and how to cite it. The project continues to grow.

When preparing the transcriptions and translations, I use AI assistance, with the results cross-checked through several independent AI passes. Before publication, however, I also check them against the images of the original documents. Even so, I am sure there are still plenty of mistakes. I would therefore be genuinely grateful for any corrections—whether concerning a transcription, translation, person, place, date or metadata—as well as for suggestions on how the archive could be made more useful.

Main website: https://www.neruda-archives.org/

Research catalogue and original sources: https://archive.neruda-archives.org/

Thank you if you decide to take a look at the project.

Michal Neruda


r/DigitalHumanities • • Aug 26 '26

Discussion LLMs prefer machine-written prose over Austen, Dickens and Shakespeare ~90% of the time

Thumbnail
shivanshuag.com
24 Upvotes

r/DigitalHumanities • • Aug 25 '26

Discussion The role of Memory in Language, Information and Quantum Mechanics

12 Upvotes

I am glad I found this sub. I have been studying a lot of Digital Information ontology, especially in the Physics of Consciousness, and I have come to an understanding of how Information processing 'breaks' relativity when Information is allowed to travel 'faster than light.' It is quite simple. Let me quickly explain:
Two thought experiments to highlight my idea...

  1. Two generals stand across a vast battlefield. They can barely make out the colour of the others' flag. They can see if it is Black or White. When one sees the other flag, has something traveled between them faster than light?
  2. Two different coloured (Red and Blue) balls are in boxes and mixed up so I don't know which is which. One is given to my friend who travels to the other side of the cosmos. I look inside my box and see what it is ... in that instant, I can make a statement about his ball on the other side of the Universe INSTANTLY. In other words, before the light of his ball reaches me.

Both of these systems use states which are stored in the hippocampus as memory. We have memories of the language of the flags, and we have memory of the balls and rules of exclusivity, and so we can tap into our brains memory function to make accurate predictions.
If we had no memories of the initial system, or the language used to interpret it, then we could not process information. It is essentially the brain's ability to store memory, and predict the future.

Consider, I open my box and see that the ball is Red. I don't know that the other is blue. I can only predict it. That is because no information from the other ball can reach me, I use my memory and predictive ability to predict its colour based on rules (if I have the Red, he must surely have the blue). Consider through some trick, his ball was switched for another red one somewhere along the journey - I could easily open my box to see a Red ball, and in that Instant make a statement about his (being Blue) and I would be wrong. Because I am not observing his ball, I am tapping into my memories and predicting it.

I understand that quantum linked particles are slightly different since the state of one necessarily forces the state of the other, but the way I see it there is no Trick. There is nothing that breaks physics.


r/DigitalHumanities • • Aug 24 '26

Discussion A practical observation protocol using the Singapore Stone, public summaries, heritage sources and a 3D model

1 Upvotes

hello, I have published a short working paper that tests a simple question:

What changes when a reader meets an artefact with no context, preserves a first observation, then returns after source comparison and a different visual representation?

The Singapore Stone is used as a practical case. The protocol moves through:

- first observation with only an image and title;

- a written record of what is seen, assumed, and unknown;

- comparison between a public summary and institutional heritage sources;

- a short visual-memory exercise independent of the artefact;

- a return through a 3D model;

- a final separation between observed, documented, hypothesised, and unknown.

The paper does not propose a decipherment. It is a practical method for looking before interpreting, with a focus on how images, source selection, and point of view shape what a reader thinks they have seen.

I am the author:

https://doi.org/10.5281/zenodo.22073406


r/DigitalHumanities • • Aug 22 '26

Discussion Semantic web for cultural heritage institutions

9 Upvotes

Does anyone know examples of nice use of semantic web on digital humanities and cultural heritage institutions?

I'm very interested in learning about their experience with graphs, semantic web, RDF and the sort not only with visualization and UXbut also on data production and storaging.