When I said the “AI” projects grifters are pushing were search engines, THIS WAS NOT THE INTENDED TAKEAWAY.
I wrote a blog post a bit ago trying to explain what the things people are presently trying to call “AI” really are, and how the whole thing is a big ol’ grift you shouldn’t touch with a ten foot pole... and I don’t think anyone actually really read that, but since I’ve written it, there sure has been a sharp uptick in stories about people you’d really hope would know better treating them like Ask Jeeves and expecting to get accurate answers to random questions. So... let me try this again.
So first let’s just cover how a search engine actually works. Or at least a simplified version of my personal understanding which is probably a bit out of date so you know, grain of salt.
While we’re used to accessing the internet through handy little URLs like, say, https://www.tumblr.com, those are just sort of handy aliases managed by this whole database setup (domain name servers) which browsers are set up to check if there’s text in there, basically, and what those match them to is numerical addresses. It’s a bit like every website has a phone number. Just as an example, open another tab and just type in oh... 74.114.154.18 and watch it bring you somewhere. They expand things a bit now and then, but basically, much like when a video game has a safe with a 3 digit combination and you don’t feel like solving a puzzle, you can totally just sit there, punch in every possible number, and doing so you’ll eventually see every website there is. It’s only what like 1,000,000,000,000 possible combinations? People who are actually serious about running a search engine will just set up a script that does that, throwing some real processing power at it, locally save every thing that comes up, and also search all those files for every file they point at, links, images, databases, whatever, and save those locally to. A whole lot of computing and a whole lot of storage later, and you literally have a local mirror of the entire internet saved to a huge pile of hard drives. Really this is such a costly endeavor it’s honestly just a handful of people who really do it and everyone else just... quietly passes your searches on to them like some kinda middleman.
Anyway, once you have your local copy of the entire internet, and entirely too much processing power to hit it with, you can do things like... look at every individual page on the entire internet and count how many times every given word appears on them, and if you’re feeling real bold, phrases, getting nice little running tallies to jam into your huge database. Then when a request comes in to give you a web page about seagull poop, you just figure well, there’s this one page that says “seagull” 49 times, “poop” 37 times, and specifically says “seagull poop” 35 times. That seems pretty on-point so you search that up as result number one.
You’ll notice there’s no thinking anywhere in here, just saving files, counting words in them, and doing some data processing on those counts. There IS a bit more to everything of course, like giving extra relevance points if something is in a title or header tag, or how somewhere along the line we all agreed to add these meta tags where people can just say outright what sort of information is on a page on the honor system (extra relevance points if people actually click links too), and someone just deciding wikipedia articles are always good so if there’s a wikipedia article, that gets a ton of bonus relevance points. Having the search string in the URL of the site of course also helps, and somewhere along the line things got gummed up with people abusing the hell out meta tags and also just giving major search engines money in exchange for bonus relevance points. Then much more recently you’ve got software engineers and suckers trying and utterly failing to “improve” results by doing dumb things, like Street Fighter 6 is out, and there’s lots of people looking for info on that, so if someone types like, “Street Fighter 3rd Strike Remy move list” into a query, well, part of that string says “Street Fighter” so let’s give all the results people searching for just that are enjoying lately, and forget the other terms.
Anyway, that’s your standard search engine. People with these sorts of “AI” projects do not, in fact, generally have a local copy of the entire internet saved. Some would like to, but you need a LOT of storage, and also there’s quite a lot of laws and security measures specifically to prevent people from doing that, and even preventing the people we as a society generally agree should be mirroring the whole internet have to leave certain parts out. Now partly they get around that by just completely ignoring that those laws exist and banking on nobody actually enforcing them in any meaningful way. Largely though they want to either avoid blatantly breaking those laws/circumventing security, so they buy “training data” from whoever’s willing to sell it, and also taking measures to obfuscate that it’s all stolen.
Anyway, you know about Markov chains? The basic idea is you have a large body of text you’ve done some statistical analysis on like we have when we archive the whole internet or what chunks we can get our hands on, and we break down the percentages of how likely every word is to come after a given word we’re looking at. Doesn’t have to be words either, you can do it with whatever. But the basic idea is, let’s say your data set is a bunch of tedious nerd posts from the year Portal came out. Now if I start off giving you the word “The” there’s all sorts of things that could come next. “The end” “the next” “the only” or maybe “the cake.” This is totally how that predicted next word thing on your phone works by the way. Anyway to really do this properly you like map out the entire web of phrases you can end up with, but for now let’s just look at that pretty popular combo of “the cake” and keep looking that way, and huh, people sure do follow “cake” with “is” these days, and especially “the cake,” I can look this up in my database easy. So you just keep hitting that next suggested word on your phone, we’re probably getting “The cake is a lie!” out of it. Someone I know loves doing stupid little things with these if you want an example.
This is totally how these “AI” things do the natural speech things, plus maybe some hard rules like “when the prompt is a single word pre-prep the chain by putting “[whatever term] is” into a standard search engine routine and just wholesale life the first sentence you can find that starts with that at the top of a block of text, then Markov chain from there.”
And we want to obscure that we’re doing this so let’s also have a rule like “OK you can go with the best match for the best work so many times in a row, but after that you have to mix it up by taking the second best word. So again, still at the height of Portal fever, we start off seeing this common word sequence, but OK let’s switch it up after “the cake is” and not go with “a” what else do we have? Well, there’s no “the” at the start, but “cake is so delicious and moist” is also real common. That’s another long string of direct quotes though, so again, let’s flag it after so many top matches and use a slightly less common one. And you end up spitting out like, “The cake is so delicious and nutritious.” Hey, that sounds like natural language, AND it’s variations on commonly said things, so it’ll probably read as legit. We’re done here, ship it.
Of course cake ISN’T nutritious, it’s like, pure sugar and gluten. But we don’t have any capacity to think or understand we’re just stringing words together based on how commonly they follow each other. Because again, there is no actual intelligence, creativity, or understanding in here, just data sets and strict procedures on how to pop words from them.
One amusing thing about this is that basically by design, it’s practically guaranteed that this is going to spit out any block of text you can imagine at you, except for the ones that are completely true and coherent. It’s intentionally avoiding ever doing that because you’d spot the plagiarizing immediately.
Oh and the whole “AI generated art” scene is doing this exact same thing. Only difference is there’s an extra step where after they download literally every image ever posted on DeviantArt, they have sweatshops full of people where for like one shiny penny a day, destitiute people pour over things, hacking them up with lasso tools and painstakingly adding meta tags for every possible thing you could for every single image they have, so the program can pull up a bunch of images that all have all the search text and then go like “OK start with this as a base, this has the 20x10 pixel blue right eye tag, does anything else in the batch have that exame tag? Cool, let’s select one of those and paste it over this eye, now, how are we on 30x40 slightly reddish upturned nose tags?” etc. etc. etc. More impressive parlor trick to pull off, but it’s still prettty plainly theft.
Anyway, this is all a thing, as I think I said, because all the people left holding the bag when everyone realized crypto/NFTs/the metaverse/etc. was a gigantic pyramid scheme have absurd amounts of processing power in big warehouses and it’s all going faulty and looking bad from being under too much heat and running too long so it’s hard to sell on eBay, so, what other scam can we do with it? Aha, fake AI.
And all the people who continue to fall for that hook line and sinker should not have the jobs they do because that level of being a gigantic mark proves them unfit to do really anything that involves any sort of decision making, do what you can to have them removed.
Also maybe give me money? I’m at risk of death otherwise.











