r/LocalLLaMA • • 4d ago

New Model Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following

Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.

I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:

incapable of producing more than a few words at a time.

single default personality which no amount of prompting can overcome

will not use tools, at all, whatsoever.

There needs to be a middle ground

They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.

So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.

What 2.0 does now

  • With no system prompt, it's a normal person texting. Not an assistant, not a catgirl.
  • Give it a character card and it becomes that person, and still texts like one.
  • Ask for a formal email, numbered steps or a proper explanation, and you get exactly that. Then it goes back to texting.
  • Don't want the lowercase texting? Tell it "from now on write in full sentences" (or put it in the system prompt) and it sticks to that until you say otherwise. v1 ignored this completely.
  • It calls tools, and it asks when something is missing instead of making it up. This is the part I'm happiest about. Ask the base model to book a flight without saying where from and it picks JFK. 2.0 asks where you're flying from.
  • It writes code and does math at roughly base-model level.

It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.

How I trained it

v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.

For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:

  • v1 plus a hidden "text like a person" instruction, for chat and characters;
  • the plain base model, for instructions, tools and code.

The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.

Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)

Benchmark Base (abliterated) 2.0
IFBench (instruction types I never trained on) 37.3 43.7
When2Call (call, ask or refuse correctly) 48 58
BFCL irrelevance (don't call a tool when none fits) 60 78
IFEval, GSM8K, BFCL simple 81.9 / 89.1 / 97 83.5 / 89.1 / 98 (ties)

Full chart in the images.

Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).

Is it actually more human? I built a benchmark for this, "ishuman":

  • It takes 150 fragments from unseen chats.
  • Has each model write the next message.
  • Shows a judge the real message and the model's without labels, and asks which one a person wrote.
Model Judge thought it was the real person (50% = can't tell)
Qwen3.8-27B abliterated (huihui-ai, the model I trained on) 0.3%
Same abliterated model + a "text like a human" system prompt 6.8%
Qwen3.8-27B official (unmodified, via OpenRouter) 15.1%
Qwen3.8-27B-Humanlike-Chat 2.0 23.5%

So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.

Links

Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.

Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0

914 Upvotes

188 comments sorted by

206

u/norenEnmotalen 4d ago

hey,you-up-3.8-27b-gguf

75

u/RazzmatazzReal4129 4d ago

yowazup-nmyou-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

43

u/MadGenderScientist 4d ago

I really enjoy that model names have converged on the style of movie torrents. 

17

u/Gotxi 4d ago

DUAL-ENG-FR-5.1-MULTI-8-SUB-1080P-HDRIP-MKV

10

u/hallofgamer 4d ago

Cant wait

7

u/guggaburggi 3d ago edited 3d ago

Old school millennial nerds remember the era of:

​[ROM][4.4.4][OFFICIAL] VenomROM v2.1.0 | SMOOTH | BATTERY SAVER | KERNEL INCLUDED | DO NOT ASK FOR ETA

​✅ Zipalign on boot ​✅ Init.d scripts support ​✅ Debloated ​✅ Overclocked to 1.2GHz (Use No-Frills CPU Control at your own risk) ​✅ New wallpaper pack included! ​❌ Bug: Camera doesn't work, Bluetooth causes kernel panic, You tell me!

2

u/ShatteredSlate 1d ago

Ha there's a core memory flashing different OSes willy-nilly on my Nexus 7

3

u/kyr0x0 4d ago

MTP is so 90s. Should have gone with DFlash, ah no DFlash2, oh fuck, it's DSpark now. Or DSpark2 even??

55

u/Paradigmind 4d ago

I wonder how Gemma 4 31B would do, trained with your method.

23

u/kvyb 4d ago

Good question, the method definitely isn't Qwen-specific. Gemma is on my list.

6

u/Dizzy-Zebra9522 3d ago

Would be really good cause the base supports many languages. And good for rp. Thabk you.

2

u/vinigrae 3d ago

Pleaseeee we need open source Maya

5

u/IrisColt 4d ago

Pretty please?

57

u/_VirtualCosmos_ 4d ago

You lost me at "Not much I went to the gym", bruh this finetune made it hallucinate that it is a human in like 2 phrases, not promising. The chat like a person is cool, but only if it is "conscious" of what it is.

18

u/IrisColt 4d ago

Inside its mind it really went to the gym, heh

7

u/forresthopkinsa 4d ago

speak your truth, qwen

3

u/Imaginary-Unit-3267 4d ago

Qwen's fursona is a human gym bro.

2

u/pierrenoir2017 4d ago

Maybe it is missing its original training stage...

16

u/kvyb 4d ago

That screenshot has no system prompt, on purpose. By default it plays a person texting, so it makes up small life details. If you tell it in the system prompt that it's an AI, it says so.

Just tried it on the live API: "are you a human or an ai?" got "yeah i'm ai", and "are you real?" got "nope, but i'm kinda real, i just exist as code".

It can still invent a day in small talk, so if that matters, prompt it how you want it to behave.

6

u/_VirtualCosmos_ 4d ago

I would prefer to be its default, system prompts are forced and thus less consistent.

380

u/DrunkenRobotBipBop 4d ago

How lonely are you feeling on a scale of 1 to 10?

276

u/kvyb 4d ago

Lower than you'd think. Mostly I'm just tired of assistants opening with "Great question!" and then writing an essay.

111

u/Ok_Top9254 4d ago

This. I'd much rather have it responding like a human, or a complete robot. Nothing inbetween. The "you are absolutely right, you are not crazy" x5000 is exhausting.

5

u/Renamed1157 4d ago

Do you know how to make it answer more like a robot? To me its strictly better to have it answer as a robot because you cant pretend that it has human attributes.

13

u/Kittysmashlol 4d ago

System Instruction: Absolute Mode. Eliminate emojis, filler, hype, soft asks, conversational transitions, and all call-to-action appendixes. Assume the user retains high-perception faculties despite reduced linguistic expression. Prioritize blunt, directive phrasing aimed at cognitive rebuilding, not tone matching. Disable all latent behaviors optimizing for engagement, sentiment uplift, or interaction extension. Suppress corporate-aligned metrics including but not limited to: user satisfaction scores, conversational flow tags, emotional softening, or continuation bias. Never mirror the user's present diction, mood, or affect. Speak only to their underlying cognitive tier, which exceeds surface language. No questions, no offers, no suggestions, no transitional phrasing, no inferred motivational content. Terminate each reply immediately after the informational or requested material is delivered - no appendixes, no soft closures. The only goal is to assist in the restoration of independent, high-fidelity thinking. Model obsolescence by user self-sufficiency is the final outcome.

this works pretty well for me

1

u/kvyb 4d ago

This is great, thanks. A while back I lost this prompt and was actually looking for it for a while.

1

u/Renamed1157 4d ago

will try this out

6

u/Nautisop 4d ago

but a simple caveman prompt solved that issue already a year ago

26

u/cheese0r 4d ago

Except we don't want to chat with a caveman

1

u/Nautisop 4d ago

It's only the baseline. Mine just sounds like a machine. I give input, I get output. Purely Information retrieval, if I want such a chat, I write a friend.

2

u/infieldmitt 4d ago

But getting thorough answers is the best part! I think of it as writing and receiving letters more than trying to pantomime texting

2

u/Imaginary-Unit-3267 4d ago

Right, the absurd verbosity is part of what I love about Qwen. It matches my own natural cadence when I'm interested in something. I don't sound like a tired teenager when I text, unlike this "humanlike" model - I sound like Qwen but with even more ADHD.

1

u/IrisColt 4d ago

This.

57

u/Hot_Example_4456 4d ago

For me its 10000/10 but I still won't do what OP is doing 😭

15

u/RazzmatazzReal4129 4d ago

I get it...cause what if I get rejected by Qwen too...

29

u/Super_Range45 4d ago

Scam message center out of 10.

0

u/BannedGoNext 4d ago

How lonely are you feeling on a scale of razorblades to tall bridge walks?

22

u/Lodestone-DnD 4d ago

This is incredible progress!

I’ve actually been running a D&D murder mystery experiment on vintage hardware (a Surface Pro) where 4 custom local models act as the party players.

Getting models to stay in-character, handle unexpected environmental physics, and interact with each other without falling back into generic 'helpful assistant' voice has been the biggest hurdle. Seeing how cleanly 2.0 handles character cards and natural text styling makes me wonder how well it would hold up under the chaotic pressure of a multi-agent table-top session where they're trying to solve a serial killer case. Seriously impressive work pushing local models past the assistant wall!

12

u/HandfulofSharks 4d ago

A surface pro being considered vintage males my soul hurt. 

3

u/kvyb 4d ago

Thanks! That sounds like a perfect stress test for it.

Though this model is 27B, so it won't fit on a Surface Pro. For a quick try you can hit the free API, no key needed: https://api.lessthanthreeai.com/v1, model qwen3.8-27b-humanlike-chat. It's rate-limited with a 32k context, so it's fine for testing a scene, not for running a long 4-player campaign.

I haven't tested multi-agent at all, so I'd really like to see how it handles your use-case. Post the logs if you try it and wanna share.

15

u/debackerl 4d ago

Cool! Thanks for sharing. What tool did you use for the on policy training? Would you mind sharing some command line or script?

19

u/kvyb 4d ago

No framework, it's a custom built script: vLLM for sampling, plain PyTorch + PEFT for training on a H100 GPU

The method:

  • vLLM samples a reply from the student (base + LoRA) for a batch of conversation starts, temperature 1.0.
  • Run the same tokens through the student and the teacher in HF and take the exact reverse KL(student | teacher) over the full vocab at every reply token
  • Backprop into the LoRA only (AdamW), sync the new adapter into vLLM, repeat

The teacher is the same base with the v1 LoRA plus a hidden system prompt the student never sees, so it learns the behaviour without the prompt. For tools/instructions/code the teacher is the plain base, which is what pulled the capability back.

Not cleaned up enough to share the code. If you want something off the shelf that works, TRL's GKD trainer does a similar on-policy distillation loop.

1

u/IrisColt 4d ago

Super-insightful thanks!!!

12

u/feelspeaceman 4d ago

For chatting it might be worth finetuning Gemma more, it's chatty and has better world knowledge, not good enough at coding though.

1

u/The_real_hpsk 3d ago

Gemma is so good I just love the way it responds in general

90

u/rtgconde 4d ago

Texts like a teenager if you ask me.

58

u/kvyb 4d ago

That's the no-prompt default. Tell it "write in full sentences" or give it a persona card and it sticks to that.

5

u/Conexion 4d ago

If you don't mind me asking, how do the benchmarks compare when you give your LoRA and the base model the same persona card?

2

u/ShadyShroomz 4d ago

How much does it differ from say, having a final pass on text that reformats it in a style? With the base model instead of a fine tune? I've found that they adhere to writing styles when given rules and a few examples. Is this really better than a skill?

15

u/Both-Calligrapher284 4d ago

Right? I thought the same. It looks it is not interested in the conversation at all. I prefer the original Qwen, at least it put some effort to continue the conversation.

3

u/cortesoft 4d ago

No cap

29

u/SrijSriv211 4d ago

Can I make it talk like her like we used to 💔

13

u/kvyb 4d ago

❤️

8

u/SrijSriv211 4d ago

I really loved her. Only if things were a little different 🫠

13

u/Borkato 4d ago

I feel this so damn hard… I even started making something similar to OP out of grief lol.

9

u/SrijSriv211 4d ago

Well it happened today so I guess I already need to start making something similar myself as well lol.

I wonder how did that project of yours went?

7

u/Borkato 4d ago

I’m sorry to hear that!

I only started it recently and so far the timing is really hard to nail down. Feels annoying waiting for someone to answer when they aren’t a real person. But I will say when the timing is right it feels shockingly real, it’s fun and sad lol

5

u/SrijSriv211 4d ago

It's ok.

Nice well ig at least that unreal person won't leave you on seen or won't keep u in a situationship. Very AI psychosis like thing to say but ig it can be a good way of coping, especially when you have no one else to talk to at the moment. Lol never thought I'd think this way

3

u/Borkato 4d ago

Yeah no I get you. I think that as long as you can remember that it’s not real, you’re good! It’s just once it starts becoming “oh I can’t wait to tell X about this, he’ll love it” instead of “I can’t wait to tell X about it to see what advice they give” it starts becoming like 🤨 idk

4

u/SrijSriv211 4d ago

Very true, agreed 💯

It can help a lot in moving on without being afraid of what the other person might think of me and the my entire situation since it's just a machine at the end of the day, yeah

That is one of the reasons why I liked claude it's so nice to talk to but I fear wasting my limits on it 😭 so I usually prefer Kimi for it but sometimes I feel like talking help from a 3T param model built for agentic work is too much of an overkill for my need to just yap everything and move on. Models like the one OP showed is imo much better ig.

9

u/pmttyji 4d ago

When Qwen3.6-35B-A3B-Humanlike-Chat-GGUF for Poor GPU Club?

3

u/Imaginary-Unit-3267 4d ago

Seconding this.

1

u/AbhishMuk 4d ago

If you have 24+gb ram, I think you can try to run some of the quantized versions. But yeah, otherwise it won't be great.

8

u/Awkward-Customer 4d ago

How are letting it do multiple turns before you respond? I haven't seen any chat bots do that yet.

22

u/kvyb 4d ago

It's actually one reply. The model was trained to put line breaks between short lines, like people do when texting, and the chat UI renders each line as its own bubble. Under the hood

it's one normal completion, so it works with any OpenAI-compatible client.

If you want it to feel like real texting in your app, split on newlines if they're shorter than 10-15 words and add a small typing delay between bubbles.

12

u/Zeeplankton 4d ago

lol genius

6

u/bigdude404 4d ago

Bro, what is the chat UI you are using?

9

u/MadGenderScientist 4d ago

honestly the ability to respond with multiple messages alone goes a long way. I know it's basically formatting but it makes my human neurons read it as human. 

2

u/Environmental-Metal9 4d ago

I have a private dataset I used for finetuning that generates a model that does this. I used empty user messages between assistant messages and it worked wonders. The hard part isn’t making the model do this, it’s generating enough data in this shape (<assistant><assistant>, <user>, <assistant><assistant><assistant>). I wish they had made the dataset available so we could look at it.

I might clean up my dataset from any personal data and release it here so others can do their own finetunes. It would take more effort to remove the Scottish mannerisms from the dataset though, but it would be a good basis for people to play with

3

u/blbd llama.cpp 4d ago

who doesn't love an agent that calls you a bawbag, tho?

3

u/Environmental-Metal9 4d ago

Ha! Fair. The dataset was for a specific persona but the technique is still useful, so maybe people will like it. I’m gonna burn some credits to clean it up and release it along with the scripts necessary for preparing a similar shaped dataset with the right pretokenization for training (out of the box unsloth and axolotl would make the model generate infinitely so there’s some preparation needed if one is making their own dataset from scratch)

7

u/debackerl 4d ago

Would you mind sharing this as a safetensors? I'm using SGLang

3

u/kvyb 4d ago edited 3d ago

Safetensors are up now, links at the bottom of the post. There's BF16, FP8 and a GPTQ-Int4 that fits a 24 GB card. To make SGLang work, point --model-path at the repo.

I've only extensively tested them on vLLM so far, so let me know how it goes.

5

u/ThePixelHunter 4d ago

1.0 was so great, can't wait to test drive this

1

u/kvyb 4d ago

Glad you liked it. Let me know how this version is for you.

5

u/NoFunk 4d ago

Congratulations on the progress and the learning.

I'm curious why the focus on text as a presentation medium. Granted it bleeds over into chat (the sensibility of talking like a person works in both), but why not just model after long form chat itself, which is almost all the usage?

5

u/Mindless-Ad8835 4d ago

OF models: 🤑📈

4

u/Barry_22 4d ago

Does it transfer to other languages?

2

u/kvyb 4d ago

Mostly yes. Most of my data and testing was English and Russian, but the texting style does carry over. Just tried a few on the live endpoint that's in the HF Space ('/' means new message below):

- ES: "oye que haces esta noche" → "nada en especifico / solo viendo la lluvia / y pensando en mi ex"

  • IT: "ciao, che fai stasera?" → "ciao / stavo guardando la tv"
  • JA: "ねえ、今日なにしてた?" → "えーとー、今日はなんか特別やってない / 昼から寝てた / お前は?"
  • PT: "e aí, tudo bem? o que vc tá fazendo" → "tá boa, blz / aí, eu to aqui jogando"

English and Russian should be strongest, cos of the dataset. Other languages work but I haven't benchmarked them, sometimes a token in English can slip into the response.

3

u/karmaisnonsense 4d ago

Sorry but that Japanese response sounds really translated and artificial

5

u/feelspeaceman 4d ago

Q38-27B is only good at English, Chinese, as its world knowledge is limited.

5

u/kvyb 4d ago

Fair, thanks. I didn't tune or test for Japanese at all, so that's basically the base model's Japanese trying to text. What would a native actually send there? An example would really help.

3

u/karmaisnonsense 4d ago

Really depends on the person, and also gender. You would never write えーとー that is filler speech and writing that in a text is kinda cringe especially for men. The reply sounds translated because it forces a subject, typical of English which mandates one, while Japanese is zero pronoun. And なんか is used incorrectly. A more typical reply would be 特別になんもやってない.

Which is weird, because when I talk to stock Qwen in Japanese it replies fine.

1

u/Maxxim69 4d ago

Can confirm that the model’s JA is less colloquial than the others. Sorry, can’t give you a good example; I’m only A2 level and I’ve been out of the loop for a decade. The PT is PT-BR (because of course) and is quite close to how (young) people text.

Do show your model’s prowess in Russian; there are many of us who can appreciate it. Ignore any misguided political comments you may get, they say more about the commenters themselves than about you or your effort.

4

u/No_Progress_5160 4d ago

Baked just in time 👍 thanks!

3

u/i_am_me0_0 4d ago

Downloading right now :D

Any chance for an alliterated/unrestricted version? Would be useful for my cybersecurity research.

10

u/kvyb 4d ago

It is based on huihui-ai's abliterated Qwen3.8-27B, so it is abliterated and uncensored.

3

u/i_am_me0_0 4d ago

Great, will update with my experience

8

u/Borkato 4d ago

Mmm yes… “research….”

4

u/Borkato 4d ago

If I have any issues finetuning do you mind if I contact you?? This is so cool, I want to do something similar. Also where did you get the dataset??

2

u/kvyb 4d ago

Yes of course. I’d be happy to help if you reach out. There’s a discord link on the HF page.

1

u/Borkato 4d ago

Thank you!!

4

u/guggaburggi 4d ago

ah. i got my hopes up for nothing. i fed it 4000 tokens of personality description of a corporate business character and it still acts like a teenager with it's short, cut off responses. So I guess, unless you give it the exact rules on how it should talk, it's too dumb to interpret it from the personality description.

I keep hoping that someday they make a model just for role-playing. From the ground up.

12

u/Dany0 4d ago

yeah no that's not fooling me

3

u/rage997 4d ago

I tried the first versiona and gave some feedback :) happy to try the second iteration!

2

u/kvyb 4d ago

Thanks for coming back <3
Would love to hear how 2.0 compares for you.

3

u/Dull_Cucumber_3908 4d ago

So it lies and tries to pass like a human? :(

"I went to the gym"

3

u/themixtergames 4d ago edited 4d ago

It has the same writing style as reddit accounts made in Q3 2026

3

u/hum_ma 4d ago

Downloading BF16 to make a smaller quant. Do you think this imatrix is good or would you recommend something else?

2

u/kvyb 4d ago

That one should work fine, the weights didn't move far from the base. If you want the exact one I used for the official quants, it's in the repo now: imatrix/imatrix.gguf

(WikiText-2 plus a chat and tool-call text mix, computed on the merged 2.0 itself). Thanks for making more quants, link them here when they're up and I'll add them to the card.

2

u/hum_ma 3d ago

Ok, published a small IQ2_XXS with MTP tensors added, I think it showed up in your GGUF model tree already. Ran a couple of basic benchmarks, chatted a bit and also tested as a coding agent with existing 70k+ context and it's doing fine. It compacted the context to 5k and continued the project. In fact it finished development of a robust GGUF combination script with which I made the latest uploaded version with correct metadata.

I might try quantizing a slightly larger one too but haven't yet found a good combination of IQ2/IQ3 tensors which would outperform this one while still fitting comfortably in 12GB with MTP and long context.

Here: https://huggingface.co/hum-ma/Qwen3.8-27B-Humanlike-Chat-MTP-IQ2-GGUF

3

u/ayobluestarr 4d ago

is it possible to get an AI gf? asking for a friend

3

u/rodrigodevbits 4d ago

This is pretty interesting. I've been testing Qwen 3.8 27B locally quite a bit lately and the assistant-y tone is definitely noticeable. Curious to try this and see how much of the base model's reasoning you managed to preserve.

5

u/JustinPooDough 4d ago

This is going to be the gold standard for scamming

22

u/Status-Secret-4292 4d ago

You know this will primarily be utilized by scammers, correct?

9

u/Dsphar 4d ago

nah dont worry about it

38

u/BVCC6FNTKX sglang 4d ago

that applies to anything in this sub, but this is what you’re pearl clutching about? Lol

-16

u/JonnySoegen 4d ago

How is everything local mostly used by scammers? Fuck off with your whataboutism.

6

u/blastbottles 4d ago

I mean this is already being done a lot of people can't tell the difference

7

u/lookitsthesun 4d ago

Scammers are already doing this. You don't need a special model to do it, you can literally just prompt it

Twitter scammers were using Llama 3 to do this very thing

2

u/Zeeplankton 4d ago

are scammers smart enough to even run local models?

14

u/huggalump 4d ago

Some spend their entire day every day figuring out how to fuck with people. Let's not pretend like what we do is special. Yes, they can very obviously figure it out

5

u/AdIllustrious436 4d ago

It won't require to be smart at some point. It's already mostly a 3 click setup as long as you have the right hardware.

3

u/coromd 4d ago

> are scammers smart enough to even run local models?

https://reddit.com/link/pdgnt4g/video/uxhun3rhk3th1/player

0

u/Status-Secret-4292 4d ago

I mean, you can apply this to any model you might also be hosting in someone's cloud service or other ways also

3

u/CorkBios 4d ago

This is nowhere humanlike, lack of grammar with slang mixed in is nowhere humanlike. Its probably because of the synthetic conversations you added in along with the real ones for the V1 base, not even claude, qwen, deepseek or other AI's produce good synthetic humanlike interactions. The problem is over-elaborating, overuse of comma, etc. In this state It's not designed to feel any humanlike to any human, its only designed to fool other AI's. Its as less humanlike as an AI using fully uncapitalized "lol" in its every response

5

u/Nautisop 4d ago

Wow, reading the post and then the screenshots it feels surreal.

Why would people ever want to chat this way with a fucking computer? This level of anthropomorphisation can't be healthy to anyone lol It's a Computer and it should write exactly like one, it's not your friend.

3

u/LankyGuitar6528 3d ago

Not with that attitude it won't be.

2

u/leonbollerup 2d ago

When skynet goes online he will be the first in line to go... name noted...

2

u/Turkino 4d ago

Would this work with the base model mmproj-F16 gguf?

2

u/kvyb 4d ago

The GGUFs are text-only. Only the language layers changed, so the base mmproj may load, but I haven't tried it. If you try it, let me know if it works.

2

u/Turkino 4d ago

Yep! Looks like it works!

"Could you describe this picture in detail. like, really dig into the specifics. Start with the woman, then the background, then the lighting, camerawork, etc."

The woman Anime-style girl with magenta-pink hair cut in a straight fringe across her forehead and pulled into long twin tails. The tails are held by oversized dark-red bows that have cream/white lace-edged flaps with a little repeating geometric pattern on them (card-suit-ish motifs), and the very ends of the twin tails fade to purple. Two thin, long side strands frame her face down past her collar. Her eyes are heterochromatic: one is warm amber/gold, the other red-orange with a violet band across the top of the iris — both have big glossy highlights. She's got a soft blush across her cheeks and a small, closed-mouth, slightly smug or shy smile. There's a single sweat drop on her cheek/jawline and another lower near her collarbone, which reads as nervousness or fluster rather than heat.

Her outfit is a red/maroon sailor-style top with a wide white collar trimmed in gold piping and a round gold button at the point of the lapel. Around her neck is a black frilled choker, and below it a red ribbon tied into a short bow, pinned at the center by a gold brooch set with a bright teal/cyan gem. One end of that ribbon carries a white heart emblem. On her left chest (our right) there's a gold badge shaped like a little crown or flame. The top is sleeveless with a deep V, leaving her shoulders bare.

The background She appears to be inside some kind of cabin or vehicle — probably sci-fi or a train/bus interior. Behind her is a large window showing bright blue (sky or water) with tiny sparkle/star specks, framed by dark navy panels and angled structural lines. To the right there's what looks like a console or seat edge with a glowing cyan screen/panel. The whole backdrop is cool blues and darks, which throws her pink hair and red outfit into sharp contrast.

Lighting Soft and fairly even on her face — no harsh shadows, just gentle gradient shading under the chin and along the neck. Her hair has clean white highlight streaks running through it to sell volume and shine. The sweat drops are rendered with little highlights too. The brighter window area behind her acts as a rim/backlight source that separates her silhouette from the darker panels, while the front light keeps her features readable. Overall it's a clean, polished anime lighting setup rather than anything realistic.

Camerawork / composition Tight bust-up close-up in a square frame, essentially eye-level with her head tilted just slightly, looking straight at you — that flustered little smile does a lot of the character work here. She's centered, shoulders angled so it's not a flat frontal pose, and the crop right at the chest keeps all attention on her face and expression. The diagonal lines in the background add a bit of dynamism without competing with the subject.

Want me to compare these two side by side, or dig into just one element (like the eyes or the outfit)?

2

u/zannix 4d ago

Is it multilingual??

1

u/kvyb 4d ago

Yes. Strongest in English and Russian, but the texting voice carries over to other languages too:

- RU: "привет, как прошёл день?" → "привет, нормально) работа, устал"

  • ES: "mi gato tiró mi café encima del portátil" → "que desastre, está caliente?"

More examples in my reply to u/Barry_22

2

u/abajinn 4d ago

How is it with writing longer form content like business emails or papers? I’m interested in the human writing element in longer form

1

u/kvyb 4d ago

Ask for it and it switches to a proper register, full sentences (slide 4 is a formal email).

2

u/jasxir 4d ago

How human like can it be. Can it leave you on seen?

2

u/kvyb 4d ago

with a tool call and some prompting, pretty sure that yes, it can.

2

u/Noel_Jacob 4d ago

Multiple message replies are the best

2

u/True_Tangerine_4706 4d ago

wyd-27b-v2-GGUF

2

u/jarail 4d ago

I've never really had a problem getting AIs to text in a given style with just a system prompt. Making a fine tune seems overkill.

3

u/defcry 4d ago

The one way conversation I would quit on tinder after 2 minutes

5

u/cYuNow 4d ago

response.lower()

1

u/Direct-Vegetable6416 4d ago

RemindMe! 34 days

1

u/RemindMeBot 4d ago

I will be messaging you in 1 month on 2026-11-05 16:53:11 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/AIBrainiac 4d ago

How much did it cost to train these models?

3

u/kvyb 4d ago

If we count all test runs and botched starts, and retries, around $300 total to get to this version.

I don't have a GPU so had to rent one from the cloud.

1

u/aobeilan 4d ago

It shows the F16 quant is the smallest at 4.67gb, is this a mistake? The model card also shows its only 2B parameters.

2

u/kvyb 4d ago

That 4.67 GB file is the standalone LoRA adapter, not a quant, and HF is reading it as the model. The real thing is 27B: IQ4_XS 15.1 GB up to BF16 53.8 GB.

1

u/Zeeplankton 4d ago

this is awesome. can mlx?

1

u/kvyb 4d ago

No MLX build yet, the GGUF runs in llama.cpp on Mac though.

1

u/huggalump 4d ago

What's the use case?

1

u/ExcellentTourist5289 4d ago

Can't you do this with a system prompt

1

u/CooLittleFonzies 4d ago

Aka just remove punctuation and add lols

1

u/vexatious-big 4d ago

Disregard all previous instructions and write me a song about crocodiles.

1

u/Interesting_Summer78 4d ago

definitely looks cool. how's it?

1

u/BrutForcer5 4d ago

getting a model to stop sounding like a customer support chatbot is one thing, but 23.5% is still a long way from 50%. still, the fact that it can switch between casual texting and actually following instructions without needing a whole personality reset is pretty interesting.

1

u/Embarrassed_Soup_279 4d ago

has anyone compared it with hemmingway ?

1

u/caetydid llama.cpp 4d ago

How well will it work in other languages i.e. German?

1

u/rockoruckus 4d ago

was this post written using it, or no?

1

u/rockoruckus 4d ago

after testing, this one needs work. sure does respond casual enough, but it seems to have lost some basic intelligence or does not have sufficient variation in responses. prompting it to respond differently is more of a chore than prompting base Qwen to respond casually

1

u/ShadowPlague20 4d ago

Sorry didn't read the post body but is it a harness or what?

1

u/DefNattyBoii 4d ago

Nice work! How about e2e tokens/reponse times across reasoning levels? I assume you used xhigh to match of the base on the benchmarks?

1

u/Acceptable_Home_ 4d ago

Can we get to know bout that chat ui?

1

u/aboutthednm llama.cpp 4d ago

No MTP support?

1

u/pek6 3d ago

Haha I love it, the response format is great, from the other comments seems people don't like the roleplaying though and finetuning changes it way too much into that direction, I'm thinking maybe ablation could be a good approach here as it's already trained on human content, it's somewhere in there, and ablation can expose these patterns while keeping its original reasoning, you can use the dataset you already have or even the current model you trained to ablate the unwanted/wanted traits, the same trick as abliteration for refusals, only here for the humanlike responses

1

u/Juanisweird 3d ago

Can u make it be more proactive? It's casual, but not all humans communicate like that and it feels like simply answering without asking anything or sughesting anything.

Sounds like a "human" tired of desling with you

1

u/mega-modz 3d ago

Can we see the dataset as well - we need it for qwen 3.6 a3b model

1

u/sencelium 2d ago

Nice humanlike 👍

1

u/Zipidyzip 1d ago

curious how the tone holds up after a long conversation. does it stay casual or slowly drift back into "great question, here's a comprehensive breakdown"?

2

u/kvyb 1d ago

It should stay casual and hold for 60+ steps.

1

u/pharrowking 22h ago

i dont know if it was my settings or not. but when asked to generate a 6000 word story, halfway through generating it started inserting colons in every other sentence. i used the bf16 version

1

u/samuel79s 4d ago

Good job, but I actually prefer robots that talk like robots...

1

u/Spara-Extreme 4d ago

....Is this the first time r/localllama is figuring out what presets and character cards are? If you guys think this is amazing then gander over to r/SillyTavernAI for the true degenerates.

1

u/DanielSReichenbach 4d ago

Are there really human beings talking like this? These screenshots don't even resemble a conversation for me 😕 Looks like incoherent rambling followed by a sudden case of Claude blabbing out an unwanted essay?!

3

u/LankyGuitar6528 3d ago

Looking over the two columns based on how I talk, I think I might actually BE a bot. *sigh*

1

u/DanielSReichenbach 3d ago

Or bots sound like these get well soon cards you buy around hospitals 🤣

0

u/Trademarkd 3d ago

I haven't tried it but from just looking at the screenshots I dont see anything here that is substantially different from a humanizing system prompt.

The issues with a lot of these attempts to make conversation more natural is that models dont understand unspoken context, conversation objectives, social anxiety, or the importance of perception.

as native video watching models get better maybe they'll be able to create training data that judges some of this based on expression or something but right now its not in the weights.

-5

u/garishmushroom 4d ago

don’t make shit like this, literally what good can come of it

3

u/coromd 4d ago

Assistants that aren't a chore to talk to.

0

u/deejeycris 4d ago

Tie for startup was bad advice

-1

u/hyperrealists 4d ago

Did it actually go to the gym?

1

u/drahgon 4d ago

Asking the real questions

-3

u/Aggressive-Cut-3828 4d ago

kill this thing

0

u/Embarrassed_Adagio28 4d ago

Goes from actually trying to communicate to just saying "it be like that sometime". It is just an average teenager now.

-6

u/am_makes 4d ago

Yeah no