It's very interesting how big of a revolution will LLMs cause (and are already causing) in genetics and bioengineering since DNA and proteins are, for all intent and purposes, really really long words
You can, very literally, fit the entire human genome in a .txt file. It's not the most efficient way to study it but you can
It's so funny that people have such strong opinions on AI but every geneticist I meet is like "oh god finally. This thing does all the work for me. We will find the catgirl gene soon"
And no this isn't just the specialized models like alphafold. You can ask Claude or Deepseek to make pipelines and bioinformatics stuff for you and they work.
A good chunk of AI and LLM research actually comes directly from bioinformatics research. To find the catgirl gene.
No technology is by default, evil.
The transformer architecture, and LLMās have very valid and very benefical advantages for humanity. At the same time, however, the raging psychopats who control these companies have indeed completely abandoned their humanity, and removed themselves from the human condition completely.
To let these people decide the fate of the world and its future is utter and sheer madness and self destruction for mankind.
I don't think they're raging pyschopaths or that they are deciding the fate of the world, let alone the destruction of humankind. I just think they want a lot of money and this stupid unequal (read, capitalist oligarch) economy is letting them get away with stupid things. LLMs are a great technology for research and yes, even entertainment. I just don't need them in every single thing I use and I don't need them to replace jobs that need a human touch.
They were using this tech to generate potential cancer-fighting proteins when I was in uni about a decade ago. I assume they can do it better now with modern models but I've been out of the field for a while. They'd use what we now call "generative AI" to come up with potential proteins, they'd be screened for the most viable (and actually makeable/refinable) candidates, and they'd be tested on cell culture assays to investigate their effect on cell growth.
Professional quibble, though I still agree with some of the sentiment. They were not using large language models to come up with potential proteins. AI is a massive field made up of a ton of different algorithms, training and mutation methods, iterative improvement methodologies, etc.
And part of the problem of the current LLM corporate led push is that it has driven research, and funding, away from alternatives to LLMs. The current generative AI that almost all of our resources are being pumped into is inefficient, generalized, and not as accurate as specialized models. It has a few things it is good at, but a lot of things it isnt, and there are far more efficient and far more effective models out there (or probably yet to be discovered) that could do even more good. But because those tools are specifically trained and specifically efficient, they arenāt widely marketable.
This document provides a comprehensive technical overview of the PairFormer module within the AlphaFold3 architecture. The PairFormer serves
^ the basis protein model, then refined by diffusion. Diffusion models are most useful for refining data down, not for generating new data altogether.
New data generation just isnāt scalable, but because these people think that all they need to do to create āartificial general intelligenceā is pump in all the knowledge in existence, suddenly every research resource is being turned towards scaling up data collection and data storage as much as possible. Regardless of if they are right, and regardless of how useful it is.















