Is anyone actually surprised by this?

you are viewing a single comment's thread
view the rest of the comments
[–] 16 points 2 years ago (3 children)

This is mildly pedantic but you're not actually running Deepseek R1, you're running a 7B version of Qwen that's been fine-tuned on Deepseek R1 outputs. All of the "distilled" models are existing models trained on R1.

  • source
  • parent
  • hideshow 3 child comments
  • [–] -3 points 2 years ago (2 children)

    Nice catch. I'll be sure after do run the real thing

  • source
  • parent
  • hideshow 2 child comments