生物技术进展 ›› 2026, Vol. 16 ›› Issue (3): 595-609.DOI: 10.19586/j.2095-2341.2025.0142

• 进展评述 • 上一篇    下一篇

Python语言在生物信息学中的应用

刘泽煜1(), 张梦杨2, 张宝宝2, 苏琳2, 李倬3, 胡文2()   

  1. 1.西北民族大学生物医学研究中心,兰州 730030
    2.甘肃警察学院,兰州 730046
    3.西北民族大学生命科学与工程学院,兰州 730030
  • 收稿日期:2025-10-14 接受日期:2026-03-10 出版日期:2026-05-25 发布日期:2026-07-14
  • 通讯作者: 胡文
  • 作者简介:刘泽煜 E-mail: 23239996400@qq.com
  • 基金资助:
    甘肃省高校教师创新基金项目(2025B-347);甘肃警察职业学院科研项目(2024GJYPTXM02)

Application of Python Program in Bioinformatics

Zeyu LIU1(), Mengyang ZHANG2, Baobao ZHANG2, Lin SU2, Zhuo LI3, Wen HU2()   

  1. 1.Biomedical Research Center,Northwest Minzu University,Lanzhou 730030,China
    2.Gansu Police College,Lanzhou 730046,China
    3.College of Life Sciences and Engineering,Northwest Minzu University,Lanzhou 730030,China
  • Received:2025-10-14 Accepted:2026-03-10 Online:2026-05-25 Published:2026-07-14
  • Contact: Wen HU

摘要:

随着高通量测序技术的快速发展,生物信息学进入多组学大数据时代,Python凭借其灵活性和丰富的工具库,已成为该领域不可或缺的分析语言。尽管Python应用广泛,但在处理超大规模数据、算法通用性及模型可解释性等方面仍面临诸多挑战,制约了其在生物信息学中的深入应用。系统综述了Python在基因组学、转录组学、蛋白质组学和代谢组学中的应用进展,重点介绍了其数据预处理、分析流程及可视化功能;详细阐述了基于Python的机器学习(如随机森林、支持向量机)和深度学习(如神经网络、图卷积网络)方法在生物标志物筛选、疾病预测及药物研发中的应用实践;并通过典型案例分析了当前存在的主要问题及改进方向。旨在为生物信息学科研人员提供Python在多组学数据分析中的系统参考,帮助解决大数据处理、模型优化等实际问题,推动生物数据分析技术的持续发展。

关键词: 生物信息学, Python, 机器学习, 神经网络, 多组学

Abstract:

With the rapid development of high-throughput sequencing technologies, bioinformatics has entered the era of multi-omics big data. Python, featured by its flexibility and abundant toolbox ecosystem, has become an indispensable analytical language in this field. Despite its wide adoption, Python still faced bottlenecks in processing ultra-large-scale datasets, algorithm generalizability and model interpretability, which restricted its in-depth application in bioinformatics. This review systematically summarized the applications of Python in genomics, transcriptomics, proteomics and metabolomics, focusing on its functions in data preprocessing, analytical workflow construction and data visualization. It elaborated on Python-based machine learning methods (e.g., random forest, support vector machine) and deep learning approaches (e.g., neural networks, graph convolutional networks) in biomarker screening, disease prediction and drug research and development. Typical research cases were analyzed to clarify current challenges and potential optimization directions. This review aimed to provide a systematic reference for researchers who adopt Python for multi-omics data analysis, help address practical issues including large-scale data processing and model optimization, and facilitate the progress of biological data analysis technologies.

Key words: bioinformatics, Python, machine learning, neural networks, multi-omics

中图分类号: