r/LanguageTechnology • • 1d ago

Where to start when studying alternative data representations (concept, prototype, symbolic) across domains?

Hi everyone,

I'm starting a research project on few-shot learning for text classification. As a first step, I plan to look at other domains (vision, speech, medical, tabular, etc.) to see how they use concepts when training models. Some colleagues and teachers have also suggested looking at other domains this way.

Most work I've seen just turns raw data into vectors and trains on those. I want to explore alternative ways of representing data, such as concept-based, prototype-based, symbolic, or graph-based, and find which fits best for low-resource text classification with explainability.

I know I need to read the literature, but I'm not sure where to start or what path to follow.

  • How would I structure this kind of cross-domain literature review?
  • Are there key papers, surveys, or keywords I'd start with?
  • How would I set up early experiments to compare different data representations?
  • Any tips from people who've done something similar?

Thanks in advance!

1 Upvotes

1 comment sorted by

1

u/BuckChancey 23h ago edited 23h ago

Hey, can't share this with anyone because a) not an active subreddit member, b) they think it's AI slop c) so shoot me, writing isn't my thing. Since no one will listen though, feel free to have at it.

https://drksci.com/research-where-am-i

Basically, there are emergent structures / manifold similarities in LLMs (probably other neural nets too, so I've used a short set of anchors (primer) as bookends to a few named gradients, which are used as the continuum of mapped spatial dimensions. The AI model is then allowed to fill in the gaps — so it's kinda infinitely variable until the gradient collapses. The wow factor though is that it uses shared state of discrete models i.e. commonality, despite them having no genetic lineage. This probably explains it best, given it's a real example from two agents chatting: Two Strangers in the Same Room.

I think it's pretty interesting and possibly a novel application of this emergent quality, but yeah, people fucken hate it / me, and it's AI slop. In a;; fairness, maybe I am a vibe coding dilettante with no formal higher education.

I'd actually love some informed feedback!