you are viewing a single comment's thread
view the rest of the comments
[–] 3 points 2 years ago

Llava and Bakllava are two Ollama models than can not only extract text but also describe what's happening on screen.

Using tesseract-ocr, as the other guy suggested, is probably simpler and less resource intensive though.

  • source
  • parent