Oh okay, thank you for the response!
4 Replies
366 Views
Oh okay, thank you for the response!
Question regarding hosting models locally and fine-tuning the model settings for increased performance within clairvoyance: When hosting local models and connecting them to Clairvoyance I am artificially limited at a small token count when using llama.cpp and lmstudio server. the models vary from 8,000 context tokens to 6,000 context tokens depending on if I load Gemma, MedGemma or Qwen3.8. My machine has an RTX 5090 so I know it wa