Seeing all posts about so and so merging with AO3 or this fandom archive moving over ...how does the process work? Like how do you avoid story duplicates or even authors? What happens when one author posts the same story under different names on AO3 and whatever old archive you're merging with? Can you filter that out? Easy of course if stories are only posted on the smaller archive and being moved to AO3 but authors post on different platforms all the time so I'm curious how the process actually works on your end.
Eskici here! I’m one of the chairs of the Open Doors committee, which is responsible for all of the imports of offline and at-risk archives to AO3.
Before each import is announced, we compile a spreadsheet with a row for every fanwork from that archive. (If the archive is backed by a database, we can often export this spreadsheet manually, but for hand-coded archives, we often do this manually.) Creators frequently email us shortly after we announce an upcoming import to let us know that their works from the archive are already on AO3 (or if they don’t want their works imported for any other reason), and we track those requests in the spreadsheet as they come in so that we don’t import duplicates.
Then, as close to right before the import as we can, we manually search AO3 for every fanwork from the archive whose creator hasn’t already contacted us. Usually, we start by entering just the title of the work and a keyword or two from the name of its fandom. We try to cast as wide a net as possible so that we don’t accidentally filter out results from anyone who didn’t tag their work with the canonical fandom tag, with tags for characters/relationships in the work, etc.
If we don’t turn up any results, we mark the work not found, but if we do find a matching work, we check to make sure the content is the same and then mark on our spreadsheet not to import it and instead invite that work to the AO3 collection that the imported archive will feed into. On the other hand, if we find too many matches to check all of them, we add keywords to the search to try to make it more manageable for our searchers.
As you can imagine, this is a lot of manual work for us to do on top of the imports themselves, which are sometimes manual as well. Open Doors recruited for import assistants earlier this year to help with tasks exactly like these ones, and we hope to do so regularly in the coming years!
If you ever have an offline or at-risk archive you’d like us to look into importing (with moderator permission), you can always reach us at [email protected]. And if you’re interested in volunteering for the project, keep an eye out for recruitment. We usually recruit for our different roles at least once or twice a year. Thanks for your interest in fanwork preservation!