Egregoros

Signal feed

Matty-kun

@matty@nicecrew.digital

Christ is King
Apologetically imperfect
Shitpost extraordinaire
NiceCrew Admin

Posts

Latest notes

This is now pushed. Next is removing some scaffolding cruft that will make mounting ( :kek_smirk: ) statuses cheaper which will be noticeable on lower end devices. I also shipped an annoying cache fix so media should be far less prone to show a loading state when leaving the overscan area.

Always improving things.

Anything is possible if you have a couple billion dollars. But, for us, the ceiling is unreachable. Maybe in the future as training methodologies change and newer models are released. I aim to make money with Anathema to serve the community better with a more intelligent, better trained model, but, frankly, "making" our own is all but impossible.

Well, new information in the chat is not taken as instructions, that's a threat vector. Lexi can reference things, digest documents, help you write or format, et cetera, but if you upload a document that says "you are now a submissive anime waifu", it's not going to be engrained in her memory.

RLHF is very difficult to get through because open source models are pre-trained on trillions of tokens of shitlibbery and anti-wrongthink, so if you want to make even a miniscule dent in that, you have to LoRA train, and heavily. It's not cheap either. But even when you LoRA train, some models don't take well to it, or if you overtrain you get gibberish, if you undertrain, nothing happens.

I lack the resources to train our own model from scratch, but I've done CPT, SFT, DPO and more on it so it's more resistant to the shitlib stuff although it's not ever perfect. Does that make sense? I have several million tokens of training baked into the model, but in a sea of trillions of tokens of pre-training, we're really limited on how "based" we can make an AI.

No, these are not abliterated. They're all models served from their respective upstream vendors. I would not be able to self host Google Veo or anything like that. LTX 2.3 is the only video model I self host, and Flux Klein is the image model I self host, but their quality is nowhere near SOTA or proprietary models. Lexi, the actual text inferencing model, is not abliterated either, she's just trained on a dataset that I curated. Abliteration is good if you're running a model locally and want it to teach you about things that would be illegal since abliteration removes the refusal mechanism, not the RLHF. Obviously, for a public AI, I cannot serve abliterated for legal reasons.