▲ 26 ▼ Alibaba Cloud releases Qwen3, an open weight LLM set that outperforms ChatGPT-o1 with only 32B parameters (qwenlm.github.io) submitted 1 year ago by gay_king_prince_charles@hexbear.net to c/technology@hexbear.net 6 comments fedilink hide all child comments
[–] JoeByeThen@hexbear.net 13 points 1 year ago (1 child) permalink fedilink source hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 9 points 1 year ago (1 child) Already on ollama. https://ollama.com/library/qwen3 permalink fedilink source parent hideshow 2 child comments replies: [–] gay_king_prince_charles@hexbear.net [S] 6 points 1 year ago (1 child) I've found Qwen preferable to DeepSeek for coding so I can't wait to try this out permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 5 points 1 year ago (1 child) I've not used Qwen yet, but I have noticed deepseek, specifically r1, is kind of a lazy coder. Lot of 'step 5 draw the rest of the owl' type responses. Unrelated, but does anyone else's internet speed come to a screeching halt when trying to download models from ollama? I swear I'm being throttled by xfinity. permalink fedilink source parent hideshow 2 child comments replies: [–] gay_king_prince_charles@hexbear.net [S] 5 points 1 year ago (1 child) That might just be LLMs in general. ChatGPT does the same. Copilot is a little more well-tuned, but I really only ever have it do boilerplate. permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 2 points 1 year ago I've had really good luck with chatgpt 4o, and, to be fair, I have teased some decent responses out of deepseek 3 (iirc). Different ways of expanding on the basic principles of asking it to 'step back and visualize different options before moving forward and fully implementing them with all necessary code, following best practices, etc.' tends to get pretty good results. permalink fedilink source parent
[–] JoeByeThen@hexbear.net 9 points 1 year ago (1 child) Already on ollama. https://ollama.com/library/qwen3 permalink fedilink source parent hideshow 2 child comments replies: [–] gay_king_prince_charles@hexbear.net [S] 6 points 1 year ago (1 child) I've found Qwen preferable to DeepSeek for coding so I can't wait to try this out permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 5 points 1 year ago (1 child) I've not used Qwen yet, but I have noticed deepseek, specifically r1, is kind of a lazy coder. Lot of 'step 5 draw the rest of the owl' type responses. Unrelated, but does anyone else's internet speed come to a screeching halt when trying to download models from ollama? I swear I'm being throttled by xfinity. permalink fedilink source parent hideshow 2 child comments replies: [–] gay_king_prince_charles@hexbear.net [S] 5 points 1 year ago (1 child) That might just be LLMs in general. ChatGPT does the same. Copilot is a little more well-tuned, but I really only ever have it do boilerplate. permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 2 points 1 year ago I've had really good luck with chatgpt 4o, and, to be fair, I have teased some decent responses out of deepseek 3 (iirc). Different ways of expanding on the basic principles of asking it to 'step back and visualize different options before moving forward and fully implementing them with all necessary code, following best practices, etc.' tends to get pretty good results. permalink fedilink source parent
[–] gay_king_prince_charles@hexbear.net [S] 6 points 1 year ago (1 child) I've found Qwen preferable to DeepSeek for coding so I can't wait to try this out permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 5 points 1 year ago (1 child) I've not used Qwen yet, but I have noticed deepseek, specifically r1, is kind of a lazy coder. Lot of 'step 5 draw the rest of the owl' type responses. Unrelated, but does anyone else's internet speed come to a screeching halt when trying to download models from ollama? I swear I'm being throttled by xfinity. permalink fedilink source parent hideshow 2 child comments replies: [–] gay_king_prince_charles@hexbear.net [S] 5 points 1 year ago (1 child) That might just be LLMs in general. ChatGPT does the same. Copilot is a little more well-tuned, but I really only ever have it do boilerplate. permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 2 points 1 year ago I've had really good luck with chatgpt 4o, and, to be fair, I have teased some decent responses out of deepseek 3 (iirc). Different ways of expanding on the basic principles of asking it to 'step back and visualize different options before moving forward and fully implementing them with all necessary code, following best practices, etc.' tends to get pretty good results. permalink fedilink source parent
[–] JoeByeThen@hexbear.net 5 points 1 year ago (1 child) I've not used Qwen yet, but I have noticed deepseek, specifically r1, is kind of a lazy coder. Lot of 'step 5 draw the rest of the owl' type responses. Unrelated, but does anyone else's internet speed come to a screeching halt when trying to download models from ollama? I swear I'm being throttled by xfinity. permalink fedilink source parent hideshow 2 child comments replies: [–] gay_king_prince_charles@hexbear.net [S] 5 points 1 year ago (1 child) That might just be LLMs in general. ChatGPT does the same. Copilot is a little more well-tuned, but I really only ever have it do boilerplate. permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 2 points 1 year ago I've had really good luck with chatgpt 4o, and, to be fair, I have teased some decent responses out of deepseek 3 (iirc). Different ways of expanding on the basic principles of asking it to 'step back and visualize different options before moving forward and fully implementing them with all necessary code, following best practices, etc.' tends to get pretty good results. permalink fedilink source parent
[–] gay_king_prince_charles@hexbear.net [S] 5 points 1 year ago (1 child) That might just be LLMs in general. ChatGPT does the same. Copilot is a little more well-tuned, but I really only ever have it do boilerplate. permalink fedilink source parent hideshow 2 child comments replies: [–] JoeByeThen@hexbear.net 2 points 1 year ago I've had really good luck with chatgpt 4o, and, to be fair, I have teased some decent responses out of deepseek 3 (iirc). Different ways of expanding on the basic principles of asking it to 'step back and visualize different options before moving forward and fully implementing them with all necessary code, following best practices, etc.' tends to get pretty good results. permalink fedilink source parent
[–] JoeByeThen@hexbear.net 2 points 1 year ago I've had really good luck with chatgpt 4o, and, to be fair, I have teased some decent responses out of deepseek 3 (iirc). Different ways of expanding on the basic principles of asking it to 'step back and visualize different options before moving forward and fully implementing them with all necessary code, following best practices, etc.' tends to get pretty good results. permalink fedilink source parent