Well, well, well. If it isn't the consequences of someone else's actions that I am directly impacted and severely affected by
let's talk about Bridgerton tea, my ask is open
Show & Tell
Jules of Nature
Fai_Ryy
No title available

Discoholic 🪩

#extradirty
macklin celebrini has autism
trying on a metaphor

Andulka

No title available
The Stonewall Inn

bliss lane
Monterey Bay Aquarium
PUT YOUR BEARD IN MY MOUTH
ojovivo
Lint Roller? I Barely Know Her
todays bird
Keni
Sade Olutola
seen from United States
seen from United States
seen from Canada
seen from United States
seen from United States
seen from Australia
seen from Myanmar (Burma)
seen from T1

seen from Canada
seen from United States
seen from Egypt

seen from China
seen from Mexico
seen from United States
seen from United States
seen from Bangladesh
seen from Argentina
seen from Australia
seen from Malaysia

seen from United States
@fandom4beginners
Well, well, well. If it isn't the consequences of someone else's actions that I am directly impacted and severely affected by
I literally love all of you, but as a Tumblr veteran, Tumblr's main feature is the reblog feature. It is the beating heart of the dashboard and the foundation for a chronological timeline. The For You page here should not be your default setting.
You guys have got to start reblogging stuff you enjoy, especially, specifically gifs and fan art but also fics and fan theories or even hot takes if you're not afraid of a lil discourse. I'm tired of being the first or third reblog for a person's post and then seeing my blog's followers do nothing but hit like, while blogs sit there with no new posts in months or years!
Reblog more stuff please. Thank you, have a good day.
You're not even going to reblog this post are you
Very true
If your business can only be reached and all info about it only be accessed via Facebook or Instagram, know that it isn’t reachable or accessible AT ALL.
The whole metaverse can no longer be properly viewed without an account and I am definitely not making one just to see your contact info or opening hours.
Get a fucking WEBSITE. It can be just a static landing page with the relevant information. But get off the metaverse!
As the French say, je suis sick of this shit. *
*can refer to many things. Currently refers to the weather being Eighty Million William Degrees outside. Gross.
I'm gonna say it, I do think that even the laziest person imaginable should have a roof over their head, food in their stomach, and access to healthcare
This is Tie, she is going to eat all of the notes
reblog to feed her notes
How is she doing this
Generative AI for Dummies
(kinda. sorta? we're talking about one type and hand-waving some specifics because this is a tumblr post but shh it's fine.)
So there’s a lot of misinformation going around on what generative AI is doing and how it works. I’d seen some of this in some fandom stuff, semi-jokingly snarked that I was going to make a post on how this stuff actually works, and then some people went “o shit, for real?”
So we’re doing this!
This post is meant to just be informative and a very basic breakdown for anyone who has no background in AI or machine learning. I did my best to simplify things and give good analogies for the stuff that’s a little more complicated, but feel free to let me know if there’s anything that needs further clarification. Also a quick disclaimer: as this was specifically inspired by some misconceptions I’d seen in regards to fandom and fanfic, this post focuses on text-based generative AI.
This post is a little long. Since it sucks to read long stuff on tumblr, I’ve broken this post up into four sections to put in new reblogs under readmores to try to make it a little more manageable. Sections 1-3 are the ‘how it works’ breakdowns (and ~4.5k words total). The final 3 sections are mostly to address some specific misconceptions that I’ve seen going around and are roughly ~1k each.
Section Breakdown: 1. Explaining tokens 2. Large Language Models 3. LLM Interfaces 4. AO3 and Generative AI [here] 5. Fic and ChatGPT [here] 6. Some Closing Notes [here] [post tag]
AO3 and Generative AI
There are unfortunately some massive misunderstandings in regards to AO3 being included in LLM training datasets. This post was semi-prompted by the ‘Knot in my name’ AO3 tag (for those of you who haven’t heard of it, it’s supposed to be a fandom anti-AI event where AO3 writers help “further pollute” AI with Omegaverse), so let’s take a moment to address AO3 in conjunction with AI. We’ll start with the biggest misconception:
1. AO3 wasn’t used to train generative AI.
Or at least not anymore than any other internet website. AO3 was not deliberately scraped to be used as LLM training data.
The AO3 moderators found traces of the Common Crawl web worm in their servers. The Common Crawl is an open data repository of raw web page data, metadata extracts and text extracts collected from 10+ years of web crawling. Its collective data is measured in petabytes. (As a note, it also only features samples of the available pages on a given domain in its datasets, because its data is freely released under fair use and this is part of how they navigate copyright.) LLM developers use it and similar web crawls like Google’s C4 to bulk up the overall amount of pre-training data.
AO3 is big to an individual user, but it’s actually a small website when it comes to the amount of data used to pre-train LLMs. It’s also just a bad candidate for training data. As a comparison example, Wikipedia is often used as high quality training data because it’s a knowledge corpus and its moderators put a lot of work into maintaining a consistent quality across its web pages. AO3 is just a repository for all fanfic -- it doesn’t have any of that quality maintenance nor any knowledge density. Just in terms of practicality, even if people could get around the copyright issues, the sheer amount of work that would go into curating and labeling AO3’s data (or even a part of it) to make it useful for the fine-tuning stages most likely outstrips any potential usage.
Speaking of copyright, AO3 is a terrible candidate for training data just based on that. Even if people (incorrectly) think fanfic doesn’t hold copyright, there are plenty of books and texts that are public domain that can be found in online libraries that make for much better training data (or rather, there is a higher consistency in quality for them that would make them more appealing than fic for people specifically targeting written story data). And for any scrapers who don’t care about legalities or copyright, they’re going to target published works instead. Meta is in fact currently getting sued for including published books from a shadow library in its training data (note, this case is not in regards to any copyrighted material that might’ve been caught in the Common Crawl data, its regarding a book repository of published books that was scraped specifically to bring in some higher quality data for the first training stage). In a similar case, there’s an anonymous group suing Microsoft, GitHub, and OpenAI for training their LLMs on open source code.
Getting back to my point, AO3 is just not desirable training data. It’s not big enough to be worth scraping for pre-training data, it’s not curated enough to be considered for high quality data, and its data comes with copyright issues to boot. If LLM creators are saying there was no active pursuit in using AO3 to train generative AI, then there was (99% likelihood) no active pursuit in using AO3 to train generative AI.
AO3 has some preventative measures against being included in future Common Crawl datasets, which may or may not work, but there’s no way to remove any previously scraped data from that data corpus. And as a note for anyone locking their AO3 fics: that might potentially help against future AO3 scrapes, but it is rather moot if you post the same fic in full to other platforms like ffn, twitter, tumblr, etc. that have zero preventative measures against data scraping.
2. A/B/O is not polluting generative AI
…I’m going to be real, I have no idea what people expected to prove by asking AI to write Omegaverse fic. At the very least, people know A/B/O fics are not exclusive to AO3, right? The genre isn’t even exclusive to fandom -- it started in fandom, sure, but it expanded to general erotica years ago. It’s all over social media. It has multiple Wikipedia pages.
More to the point though, omegaverse would only be “polluting” AI if LLMs were spewing omegaverse concepts unprompted or like…associated knots with dicks more than rope or something. But people asking AI to write omegaverse and AI then writing omegaverse for them is just AI giving people exactly what they asked for. And…I hate to point this out, but LLMs writing for a niche the LLM trainers didn’t deliberately train the LLMs on is generally considered to be a good thing to the people who develop LLMs. The capability to fill niches developers didn’t even know existed increases LLMs’ marketability. If I were a betting man, what fandom probably saw as a GOTCHA moment, AI people probably saw as a good sign of LLMs’ future potential.
3. Individuals cannot affect LLM training datasets.
So back to the fandom event, with the stated goal of sabotaging AI scrapers via omegaverse fic.
…It’s not going to do anything.
Let’s add some numbers to this to help put things into perspective:
LLaMA’s 65 billion parameter model was trained on 1.4 trillion tokens. Of that 1.4 trillion tokens, about 67% of the training data was from the Common Crawl (roughly ~3 terabytes of data).
3 terabytes is 3,000,000,000 kilobytes.
That’s 3 billion kilobytes.
According to a news article I saw, there has been ~450k words total published for this campaign (*this was while it was going on, that number has probably changed, but you’re about to see why that still doesn’t matter). So, roughly speaking, ~450k of text is ~1012 KB (I’m going off the document size of a plain text doc for a fic whose word count is ~440k).
So 1,012 out of 3,000,000,000.
Aka 0.000034%.
And that 0.000034% of 3 billion kilobytes is only 2/3s of the data for the first stage of training.
And not to beat a dead horse, but 0.000034% is still grossly overestimating the potential impact of posting A/B/O fic. Remember, only parts of AO3 would get scraped for Common Crawl datasets. Which are also huge! The October 2022 Common Crawl dataset is 380 tebibytes. The April 2021 dataset is 320 tebibytes. The 3 terabytes of Common Crawl data used to train LLaMA was randomly selected data that totaled to less than 1% of one full dataset. Not to mention, LLaMA’s training dataset is currently on the (much) larger size as compared to most LLM training datasets.
I also feel the need to point out again that AO3 is trying to prevent any Common Crawl scraping in the future, which would include protection for these new stories (several of which are also locked!).
Omegaverse just isn’t going to do anything to AI. Individual fics are going to do even less. Even if all of AO3 suddenly became omegaverse, it’s just not prominent enough to influence anything in regards to LLMs. You cannot affect training datasets in any meaningful way doing this. And while this might seem really disappointing, this is actually a good thing.
Remember that anything an individual can do to LLMs, the person you hate most can do the same. If it were possible for fandom to corrupt AI with omegaverse, fascists, bigots, and just straight up internet trolls could pollute it with hate speech and worse. AI already carries a lot of biases even while developers are actively trying to flatten that out, it’s good that organized groups can’t corrupt that deliberately.
alright everyone is being sassy but nobody has brought up the actual reason why scientists are interested in the titanic in the notes
Per wikipedia:
The Titanic was made of steel, presumably to resist corrosion (i mean. thats why boats are made of steel i assume) but even when iron/steel rusts people did not expect to find like. decomposition. bacteria are EATING the titanic.
and there's wood furniture from the titanic that isn't decaying. hell, there are wood ships at the bottom of the ocean that archaeologists study. so the expectation people have for the titanic is not that the steel would decompose at the bottom of the ocean. Even in 100 years, since there are much more ancient preserved wooden ships iirc.
(im not particularly knowledgeable about ships, i just had heard about the science going on around the titanic so i wanted to clarify that on this post for people)
The Titanic is an exceptionally weird whalefall basically.
Not a marine biologist but biologist enough to weigh in on this. The reason we have iron eating bacteria but not wood eating bacteria at the bottom of the ocean is simple. Hydrothermal vents release a cocktail of different mineral ores into the ocean. And bacteria and other organisms evolved alongside those so they evolved to digest these, like iron or other metal ores
Wood however does not exist at the bottom of the ocean since it basically never sinks down, even when logs are flushed out into the ocean they basically never end up at the ocean floor. So there's no bacteria that evolved to decompose lignin, which is already complex enough to decompose on the surface. And that's why wooden ships or the furniture on the Titanic stay intact for hundreds of years or longer.
Okay, but I remember reading a magazine stating that the titanic would basically be gone by now because of the rust eating the ship.
The sister ship HMHS Britannic is still perfectly preserved because of the way it sank, and partly in due to the coral that has grown around it keeping the structure in tact I believe was the reason.
Coral acting like a living fossil specifically for sunken ships is the coolest thing I've ever heard
So lemme get this straight, the Titanic is a whalefall, the Britannic is a biological mummy, and all the wood ships that sink into the deep ocean are preserved incorrupt because nothing can eat them.
Zombie, mummy, lich, respectively.
when tumblr dies i'll live under your bed and you can say out loud what you would post and i will say LIKE or REBLOG it'll be just like we're still here
Shitty comic day i didnt forget you this year
Reblogs r back on so BEHAVE.
“packmate” is a very good word to describe how i feel about some people. like we’re more than friends in my eyes but it’s not romantic you’re just. my person and im #keeping you. because you’re in my pack. wags my tail
In the club
I think I’m literally never gonna be sick of this masterpiece. I think watching it on a loop for eight hours could fix me. Dancing’s what clears my soul. Dancing’s what makes me whole.
I just love that this very video is an accumulation of thousands of years worth of art made by people who have never met each other. The concept of this video was so completely unfathomable to every single artist who made the sculptures and yet they’ve all put something toward the creation of it.
ITS BACK ON MY TIMELINE
ai fatigue
ai books, fanfics, ai google assist, ai short form content, ai summaries, ai memes, ai artists
I cant even go and see some street food trucks without their damn menus having AI images on it. you have a camera. take photos of your goddamn food. I would rather see a photo of the most weirdly lit burger than the fake monstrosity you're showing me.
I am cordially inviting all of you to post this in the comments/reblogs of any and every Al post you come across, on any platform. Especially useful for accounts who “teach” people how to use Al.
I hate it when you’re reading smut and you can’t figure out what position they’re in.
sometimes it just ends up being something like
ITS BACK
Y’ALL NEED JESUS
Please stop reblogging this post
This post made my water break
In honor of my daughter’s first birthday next week, I’m sharing the post that made me laugh so hard that it broke my water.
WHAT
God, I love this accursed website.
Hey internet, the girl that was born from this post is 4 years old today (July 2 2021) also, the gif still makes me laugh. Happy Birthday, Marceline!!
Happy July 2nd, 2024, guess Marceline is 7 now
Happy July 2nd, 2025; she’s 8 now (5am for me)
Happy Birthday to my Marceline, 8 years ago today (July 2, 2025) she came into this world five days before her due date to see wtf was so funny 🤣 this post, like many other things, are on a long list of topics to discuss when she’s older lol
Happy 9th Birthday Little Marceline!!! (July 2 2026)
It’s July 2 2026 internet, help me wish a Happy Birthday to my brilliant and beautiful 9 year old, Marceline! I’ve shown her (most of) the comments and she says “thank u internet 😁😁”