For drug discovery, protein-folding models such as AlphaFold have a data problem: there isn't enough of it in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue. Protein structures' locked away by the thousands in drug company vaults' offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly. The group used OpenFold3 ' an open-source replication of AlphaFold 3 ' to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparable ones trained on public data alone and those trained on the siloed datasets of individual firms. The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available. The findings, he says, strengthen the case...
learn more