Self hosting an LLM for research

Maroon@lemmy.world · edit-2 4 months ago

Self hosting an LLM for research

Terrasque · edit-2 6 months ago

Reasonable smart… that works preferably be a 70b model, but maybe phi3-14b or llama3 8b could work. They’re rather impressive for their size.

For just the model, if one of the small ones work, you probably need 6+ gb VRAM. If 70b you need roughly 40gb.

And then for the context. Most models are optimized for around 4k to 8k tokens. One word is roughly 3-4 tokens. The VRAM needed for the context varies a bit, but is not trivial. For 4k I’d say right half a gig to a gig of VRAM.

As you go higher context size the VRAM requirement for that start to eclipse the model VRAM cost, and you will need specialized models to handle that big context without going off the rails.

So no, you’re not loading all the notes directly, and you won’t have a smart model.

For your hardware and use case… try phi3-mini with a RAG system as a start.