Xinjiang University — Mathematics & Applied Mathematics (National Base), Urumqi Tsinghua TEEP · Shenzhen X-Institute — Joint Program · Zero One Scholar
China College Student Self-improvement Star (2025)
IEEE Biometrics Council · Chinese Chemical Society (CCS), Student Member
· ·
个人简介
王一权是新疆大学数学与应用数学专业(国家理科基地班)的本科生,同时也是清华大学钱学森班与深圳零一学院的联合培养零一学者。他的研究方向为科学智能(AI for Science),致力于解码生命系统的复杂性。
他的科研之路始于数学建模。本科早期,他在建模竞赛与科研项目中广泛探索,从图论到 AI 驱动的热浪灾害风险分析,论文列表里的跨领域条目正是那段旅程留下的脚印。数学建模训练了他把真实系统抽象为模型的直觉,也让他逐渐看清了自己真正的热情所在。蛋白质是研究生命系统的合适小模型,Anfinsen 法则指出氨基酸序列决定三维结构,问题定义清晰,序列、结构与功能注释数据完备,天然适合表征学习,加之个人兴趣,他将蛋白质确立为核心研究对象。如果说蛋白质是理解生命系统的小模型,基因组就是这套系统更完整的蓝图,博士生阶段他的研究将聚焦基因组,同时继续推进蛋白质组的工作。
研究成果发表于《Journal of Chemical Information and Modeling》《Physica D: Nonlinear Phenomena》《Information Sciences》等期刊,在IEEE BIBM、IEEE SMC等计算机主会上发表,并在ICLR、ICML、AAAI等会议的研讨会(workshop)中展示。欢迎就共同感兴趣的课题交流讨论。
Wang, Y., Cai, M., et al. (2026). Sequence context decodes multi-state conformational heterogeneity from crystallographic B-factors. Science China Life Sciences, in press (earlier version: LMRL Workshop at ICLR 2026).
即使晶格堆积掩盖了部分信号,静态晶体结构中仍保留着蛋白质运动的线索。BFC 利用蛋白质语言模型提供的序列上下文,从 B 因子中恢复逐残基的构象异质性,与晶体学结构系综的异质性分布达到 0.83 的 Pearson 相关系数。分子对接、分子伴侣和抗体案例展示了校正信号在结构分析中的用途。
Yiquan Wang is an undergraduate student in Mathematics and Applied Mathematics at the National Base for Research and Teaching Talents at Xinjiang University, and a joint-training scholar in the Tsien Excellence in Engineering Program at Tsinghua University & Shenzhen X-Institute. His research focuses on AI for Science, driven by an enduring pursuit to decode the complexity of living systems.
His research journey began with mathematical modeling. In his early undergraduate years, he explored widely through modeling competitions and research training programs, working on graph theory and AI-driven heatwave risk analysis, directions that left their footprints across the interdisciplinary entries in his publication list. Mathematical modeling trained him to abstract real-world systems into models, and gradually revealed where his true passion lay. Proteins are a suitable small model for studying living systems. Anfinsen's dogma states that the amino acid sequence determines the three-dimensional structure, which makes the problem well-posed, and sequence, structure, and functional annotation data are abundantly available, a natural fit for representation learning; together with his personal interest, this led him to settle on proteins as his core research object. If proteins are the small model for understanding living systems, the genome is the blueprint of the same system in full. In his PhD stage, his research will focus on genomics while continuing his work on proteomics.
He has since established two distinct yet complementary research directions. The first centers on data-driven Representation Learning and Biomolecular Mechanisms, applying deep learning to a broad range of life science problems spanning protein structure prediction, protein function prediction, protein design, molecular dynamics, multi-omics, and molecular mechanisms. The second pursues First-Principles of Biological Complexity Based on Mathematical Physics, deriving analytical ground truths of biological systems through condensed matter physics, topology, manifold learning, and differential geometry, independent of AI approximations. His overarching goal is to synergize these two directions, combining the inductive biases of physics with the expressive power of deep learning to tackle a constellation of formidable open problems, including Intrinsically Disordered Protein (IDP) dynamic ensemble generation, Liquid-Liquid Phase Separation (LLPS) and 3D genome folding, rational design for "undruggable" targets, and the construction of Virtual Cells. Through these advances, he ultimately aims to confront complex human diseases and move closer to a fundamental understanding of life itself.
His work has been published in journals including the Journal of Chemical Information and Modeling, Physica D: Nonlinear Phenomena, and Information Sciences, presented at main conferences including IEEE BIBM and IEEE SMC, and presented at workshops of ICLR, ICML, and AAAI. Discussions and collaborations on topics of shared interest are always welcome.
Main Courses: Mathematical Analysis, Advanced Algebra, Analytical Geometry, Partial Differential Equations, Functional Analysis, etc.
Program: National Base for Research and Teaching Talents in Basic Sciences "Mathematics". Mathematics Base Introduction
Yan-Hong Qin: Then at the College of Mathematics and System Sciences / Institute of Mathematics and Physics; now at the School of Physical Science and Technology.
Kai Wei: Xinjiang Key Laboratory of Biological Resources and Genetic Engineering, State Key Laboratory Incubation Base co-built by the Province and Ministry, College of Life Science and Technology.
Research Interests: Computational biophysics, including IDP dynamic ensemble generation; decoding LLPS & 3D genome folding; rational design for "undruggable" targets.
Research Experience
Featured Papers
Direction 1: Representation Learning & Biomolecular Mechanisms
My current research leverages state-of-the-art Deep Learning techniques to decipher complex biological systems. I focus on overcoming the limitations of traditional bioinformatics by introducing novel inductive biases and representation strategies. My work spans protein structure prediction, protein design, and molecular dynamics, aiming to provide robust, high-throughput tools for the life sciences.
Wang, Y., Cai, M., et al. (2026). Sequence context decodes multi-state conformational heterogeneity from crystallographic B-factors. Science China Life Sciences, in press (earlier version: LMRL Workshop at ICLR 2026).
Static crystal structures retain clues to protein motion even when lattice packing obscures them. BFC uses protein-language-model sequence context to recover residue-wise conformational heterogeneity from B-factors, matching crystallographic ensemble profiles with a Pearson correlation of 0.83. Docking, chaperone, and antibody case studies illustrate how the corrected signal can support structural analysis.
Amino acid sequences can be read as musical scores: hydrophobicity sets pitch and molecular weight sets rhythm. Converting these sequences into two-dimensional spectrograms provides useful representations for protein function prediction. Benchmarking and ablations show that much of the predictive signal comes from the 1D-to-2D transformation itself, with biophysical encoding providing a further gain; a computational proof-of-concept extends the representation to GFP design.
Direction 2: First-Principles of Biological Complexity Based on Mathematical Physics
While Deep Learning provides powerful approximations, my ultimate goal is to find the analytical ground truths of biological systems. I am transitioning from purely data-driven approaches to physics-informed theories. I aim to use Condensed Matter Physics, Topology, Manifold Learning, and Differential Geometry to describe the energy landscapes of proteins and the thermodynamic essence of life.
An exact decomposition of the Hasimoto–DNLS effective potential separates chirality from the contributions of local backbone geometry. Analysis across 856 proteins shows why this compact representation describes folded structures yet fails to predict native folds through a local real-potential reduction. The map is a kinematic identity, and its dispersion residual provides a geometric marker of near-integrable α-helices.
Helices and coils leave distinct spectral signatures in backbone geometry. Mapping 1,986 protein structures through the discrete Hasimoto transform reveals abrupt, directionally asymmetric boundaries between low-entropy helices and broadband coils. Pointwise integrability and windowed spectral probes capture complementary features; the analysis also explains boundary assignment ambiguity and the spatial–spectral trade-off that limits windowed measurements.
Direction 3: Interdisciplinary Explorations & Collaborative Synergies
My research journey is fueled by broad curiosity. I have collaborated with experts in linguistics, computer vision, cryptography, and operations research. These diverse experiences have equipped me with a unique toolkit—allowing me to transfer methodologies and tackle problems from orthogonal perspectives.
Clustered failures challenge network reliability differently from isolated faults. The Region-Based Fault model represents damage as spatially constrained connected subgraphs and establishes conditions for Hamiltonian connectivity in k-ary n-cubes with odd k ≥ 3 and n ≥ 2. A constructive adaptive algorithm supports the proof, while experiments identify separation between fault clusters as a key factor in resilience.
Heatwave risks propagate across disciplinary boundaries, but the evidence is scattered across thousands of studies. HeDA organizes 8,365 publications into a knowledge graph with 34,933 entities and 42,890 directed relationships. Held-out reasoning tests show benefits from graph augmentation, while the reconstructed topology highlights agriculture and human health as structural mediators linking thermal stress to wider socioeconomic losses.
The full record of research projects, study programs, internships, competitions, and reviewing lives on the Experience page.
WeChat Official Account
biomath
Exploring the frontiers of AI for Science and Computational Biology: from protein design, molecular and macroevolution, to multi-scale physicochemical simulations.