Looping was a problem after reaching a certain context window size. The llama.cpp flags - -flash-attn on and looping penalties helped.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: