You can run model locally, a gaming PC is enough, may-be not for cuting edge models but if you want to generate clues for a RPG or rephrase a letter it's good enough. You can look for LM studio and Stability matrix for example
post
It is absolutely possible to run 100% locally, but in practice, at the high end, only quantized models. The full size top end models require hundreds of GB of video ram, and while you can buy that, it's stupidly expensive. Quantized models can often perform nearly as well with a small fraction of the ram, but they do sacrifice a little in precision.
These models (well the good ones) ultimately all trace their origins to what you'd likely consider "stolen" data. Whether that's ethical is debatable. If you're in the "information should be free" camp, there may not be an issue here.
As for the power/environmental impact, for what they do LLMs are actually very low impact per-request. If you're concerned about your personal AI power use, then I hope you never fly in an airplane, and minimize your driving because those are much bigger issues.
It's the scale of use that makes AI an environmental problem, and that's a question about corporate use of AI, not personal use of AI.
The models can be ran locally, yes. I’m doing it.
As for the ethics of training? It’s not unethical and I’ll debate that all day, if you want to get into it.
No one ever does though, they just rage quit once I get them to a certain point🤷♂️.
From that perspective local LLMs sound more like classical piracy. Not ethical, not FOSS, but out of the hands of greedy corporates.
Some companies build out more electrical capacity than they use. Substations are funded by the DC and built to twice the desired capacity - and half is delivered to the community. The generators are run only when utility power goes down, and run on diesel.
However, not all companies are this responsible and the grievously irresponsible ones get the press and make the responsible companies look bad.
I'll take a crack at it.
I run AI on an old Nvidia p40 purchased online for less than $200. The models I run on it came predominantly from Chinese companies and groups. On balance, they:
-
use much less fossil fuel in producing these LLMs than their Western counterparts
-
produce models that run much better on lower end consumer hardware
As for training data, the inputs used to create these models, it varies greatly but the most popular line, Qwen, Came from the company's own data from operating such huge networks and systems for so long
I really fail to see how any of that is worse than playing a video game.
None of this takes away from the very real issues around data center build out and Western companies using the systems to scare people and continue an economic bubble. That's all true and bad. But there's nuance. Not all AI is created equal
They are all trained on copyrighted material without permission, no LLM is ethical
do you consider "piracy" like zlibrary and annas archive unethical?
Human brains are all trained on copyrighted material without permission, no human brain is ethical
You can also train your own local models with license free material if you wish! I think one of the easiest ways to get into that is by using software like unsloth (that's the one i am using), an open source no-code tool which can be both used to train models on whatever data you wish and to run models either locally or using an inference provider.
Quick example for something like that which is also not dependent on copyrighted material is RAG, where you can provide the 400-page manual for something and then can chat with "the document" to get explanations, ask quick questions without searching for possibly multiple occurrences of a specific term and similar stuff.
Some don't speak any "language" they are trained to "speak" and "think" in terms of election orbitals and bonding energy. They are used in pharma and materials science to work on intractable problems like superconductivity and meds for Parkinson's.
I run models locally on my 5070 i bought a year back for around 550€ (now costs around 900€ - insanity). Yes, it runs completely locally with pretty good results, and i am using an AM4 platform, which is now a few years old, and was also able to run smaller models with my 3070Ti with good speed.
My viewpoint regarding the models themselves is that since everything on the web has been used including everything i have put on the web in the last decades, there is also not much of an ethical problem here when using these things in a non-commercial setting - i am not making any money with it and i am not disseminating the output, and therefor i also do not cut into the profits of any creator (i mainly use it to automate tedious tasks or to get a complicated RegEx/sed/awk situation under control without breaking my brain - no cultural output like text or images). On a larger scale i would prefer it if models and datasets were under UNESCO stewardship - free for non-commercial use, with licenses being sold to corporations and organizations, where the income from the licenses get distributed to the people providing for the datasets (in this case in an opt-out basis - these models already represent a cultural snapshot of humanity, and the assumption is that artists WANT to add to human culture) or for financing young artists.
Not an expert but I've read up a little.
One can pay to license material for LLM's (Google at. al. are not fans of this ofc.), or simply only use stuff in the public domain or your own internal material.
Private/local models exist for various use cases like not wanting even the chance of your stuff hitting the cloud. Ex. https://ollama.com/
If you have a 4070 or so you can run a smaller model without a issue, which is well within hobbyist range, but larger models quickly balloon hardware requirements to the point where you're putting in some bucks.
They call it open-weighted and not open-source because they don't have the training material (stolen books lol) but still want to pretend they are l33t hackers not bound to BigTech.
Yes, it's possible to use it locally with a $2000 computer (that's for the cheapest ones) but you'll only get a few words per second. It doesn't matter to the vibe coders who don't know how to code.
The software to run that is open-source but the training of the model (the "open-weight" black box) requires to destroy the environment at least once.
Last but not least the free models are obviously censored but people don't care about censorship anymore for some reason. "Tiananmen didn't happen? Not my problem" without understanding that more is hidden.
Anyway, no, there is nothing ethical about it.
a) Yes, today it costs 2000$ because of the cost explosion - my setup cost me around 1400$, and my GPU was already bought when prices were rising. No, i do not get a few words per second, i get around 50-60 token/s, which is more than enough for personal use. (Edit: and that is WITH CPU offloading, where layers that don't fit into VRAM get placed into system RAM, and i am still running DDR4 to boot)
b) If the completely insane AI corpos would stop training humongous models to chase after non-achievable AGI, the one-time investments would have paid off by now for local use. This situation has nothing to do with local models but insane billionaires and execs.
c) go ahead and lookup abliterated (not a typo) models on HuggingFace - these do NOT refuse any requests, because that has been pruned out. These answer everything about Tiananmen, DEI topics like erosion of LGBT and womens rights in the west and whatever atrocities any group might have commited, while also telling you about whatever you want to know. I run only these models, because i refuse to partake in censorship.
Edit: Regarding the training material: This training material also consists of MY output over the decades on the web. Therefor, i do not have qualms using these models for non-commercial usage without disseminating the output - classic personal usage, mainly for automating tasks that are not easily done by hand such as grabbing game descriptions from steam, condensing them down to a short sentence and putting the result in a database, together with the corresponding steam tags, and automatically searching the web for information about the game when there is no steam page for it - which is an insane amount of work to do by hand for my library of more than 30k titles. I let it work during off hours on that task.
AI, yes. Generative AI, no.
Example of ethical use: in addition to human analysis, you can use AI to analyse mammograms. It can then peruse a huge database of images and diagnoses and may find tumors that a human alone may not have spotted.
"Open weighted" I take to mean that the user can see what sources the LLM/AI was trained upon. Usually this information is a closely guarded corporate secret.
As for "models" this is no different than the different models that a car company offers. Sentra from Nissan, F-150 from Ford, Civic from Honda
If locally run or through a data center, LLM/AIs have two major problems:
- The results can't really be trusted due to their tendency to hallucinate.
- Frequent LLM/AI usage has been proven to deskill the user as well as a reduction in cognitive ability. I'm dumb enough without getting help from a clanker.
Assuming "run it 100% locally" means on your own hardware with network access, I see no reason this can't happen given a powerful enough computer.
all 21 comments