HAddiT: Accurate Hydrogen Addition, Reliable Simulations

HAddiT (Hydrogen Addition Tool) is built to solve this challenge. It delivers accurate, chemically consistent, and quick hydrogen addition to the molecular structures, thereby enabling efficient and dependable high-throughput Molecular Simulations.

1. Introducing HAddiT

Computational chemistry and molecular modeling using physics-based simulations have become indispensable for elucidating and predicting chemical and biological phenomena. The rapid growth of high-resolution experimental structures, spanning small molecules in solid-state and protein–ligand complexes, has strengthened the reliability of these in silico workflows. However, a persistent bottleneck remains: accurate addition of hydrogen atoms during the system preparation stage, as they are often missing or ambiguous in experimental structures. Incorrect filling of heavy atom valencies with hydrogens can propagate critical errors across Molecular Dynamics, Quantum Mechanics, Free Energy Perturbation, and structure-based drug discovery pipelines.

HAddiT (Hydrogen Addition Tool) is built to solve this challenge. It delivers accurate, chemically consistent, and quick hydrogen addition to the molecular structures, thereby enabling efficient and dependable high-throughput Molecular Simulations.

2. Why is Hydrogen Addition to Heavy Atoms Important in Computational Chemistry?

Experimental structure-determination techniques such as X-ray crystallography, NMR, and Cryo-EM often fail to provide reliable hydrogen atom coordinates due to their low electron density and poor scattering. Consequently, most experimentally resolved molecular structures lack explicit hydrogens. However, physics-based modeling methods such as Molecular Dynamics (MD) and Quantum Mechanical (QM) calculations require a complete hydrogen representation. Computational chemists must therefore add hydrogens to generate simulation-ready starting structures. Accurate hydrogen addition is essential for obtaining physically meaningful and reliable modeling results.

3. Why Do Existing Methods Fail?

Conventional hydrogen-addition tools such as OpenBabel, RDKit, PyMol, Avogadro, AutoDockTools, etc., rely on heuristic bond-order perception and fixed valence rules. These methods infer bond order from the bond lengths in the input structure and the known equilibrium bond distance for the constituent atom types. This approach tends to break down when the molecular structures deviate from equilibrium geometry.

Crystal structures that deviate from equilibrium geometries, protein–ligand interactions, solvent-exposed environments, and flexible or heterocyclic molecules frequently encounter misassigned bond orders due to the deviation of the bond length from the known equilibrium value. The result is incorrect valency estimation, leading to the addition of too many or too few hydrogen atoms. These errors propagate into downstream simulations and limit the ability to build fully automated, reliable high-throughput pipelines for Physics-based modeling.

4. Our Solution: What HAddiT Provides

HAddiT delivers chemically reliable hydrogen addition where conventional tools fail. It provides:

  • Accurate hydrogen addition across complex, heterocyclic, aromatic, and charged molecules
  • Verified chemical consistency of the hydrogen-added structures
  • Automation-ready output for simulation pipelines

By ensuring structurally correct, hydrogen-complete models, HAddiT removes a critical bottleneck in molecular modeling workflows.

5. Case Study

Note – The assignment of the protonation states of titratable sites in a molecule based on their pKa and the solution pH is not included in this benchmark study.

i) Precision Modeling for PROTAC Modalities

Proteolysis-Targeting Chimeras (PROTACs) have emerged as a transformative drug discovery modality in recent times, with late-stage clinical candidates such as Vepdegestrant (ARV-471) and Bavdegalutamide (ARV-110) demonstrating strong therapeutic promise. Understanding PROTAC efficacy requires accurate modeling of protein of interest (POI)–E3 ligase–PROTAC ternary complexes, where stability and cooperativity dictate degradation efficiency. Physics-based approaches such as molecular dynamics (MD) simulations are central to these insights but are critically dependent on chemically correct starting structures.

To assess preprocessing robustness, we evaluated all the PROTAC ternary complexes available in the RCSB Protein Data Bank using conventional hydrogen addition tools, including Open Babel and RDKit. These methods failed in ~25% of cases due to bond-order and valency errors, leading to incorrect hydrogen addition and unreliable MD/FEP inputs. HAddiT resolved these inconsistencies w.r.t. the SMILES or 2D structures, delivering chemically consistent, simulation-ready PROTAC structures that enable robust MD workflows, FEP-based linker optimization, and confident ternary complex analysis.

ii) Precision Modeling for Small Molecules

To evaluate HAddiT’s performance across diverse scenarios, we tested the tool on small molecules in their solid-state crystal structures and in their protein-bound states. In each case, HAddiT’s performance was benchmarked against hydrogen addition using conventional tools.

  1. From Experimental Crystal Structures to Simulation-Ready Solid Forms

The Cambridge Crystallographic Data Centre (CCDC) hosts experimentally determined crystal structures for over 2,500 drug molecules, spanning multiple polymorphs and solid forms. These structures are critical for solid-state modeling, lattice energy calculations, in silico crystal structure prediction (CSP), and AI-driven molecular design.

However, hydrogen atoms in X-ray crystal structures are usually unresolved. We observed such failures in clinically important drugs, including azathioprine (an immunosuppressant used to prevent renal transplant rejection, treat rheumatoid arthritis, Crohn’s disease, and ulcerative colitis) and aminocaproic acid (an antifibrinolytic agent used to induce clotting postoperatively), where standard hydrogen addition tools produced chemically inconsistent models. HAddiT resolved these limitations by restoring chemically accurate bond orders and valence states within complex crystal environments. The result is simulation-ready crystal structures for physics-based solid-state modeling workflows and CSP pipelines towards CMC applications.

  • 2. Ensuring Chemical Fidelity in Protein-Bound Ligand Modeling

Accurate protein–ligand representations are fundamental to binding free-energy calculations and AI-driven drug discovery. Analysis of over 461,000 BioLiP entries revealed a high-confidence subset of ~1,700 unique drug-like molecules positioned within protein binding pockets. However, because hydrogen coordinates are rarely resolved experimentally, researchers depend on automated protonation tools that often introduce chemical inconsistencies at scale, undermining the reliability of downstream MD and FEP simulations.

The importance of precise hydrogen addition is exemplified by Mibefradil (DB01388), a T-type calcium channel blocker withdrawn after being identified as a mechanism-based (suicide) inhibitor of CYP3A4, leading to severe drug–drug interactions. Structural analysis of the CYP3A4–Mibefradil complex (PDB ID: 6OO9) shows that conventional hydrogen addition tools fail to produce chemically accurate structures in the bound state. HAddiT corrects bond order, formal charge, and stereochemical errors, restoring accurate hydrogen placement and delivering simulation-ready structures critical for metabolism-aware modeling and drug safety assessment.

  • 3. Docking-Derived Enzyme–Substrate Complexes for Metabolic and ADME Modeling

High-throughput virtual screening and enzyme engineering pipelines depend on physics-based molecular docking to generate protein–ligand and enzyme–substrate complexes. These docked structures are routinely advanced to MD-based binding free-energy calculations, hit-to-lead optimization, and metabolic pathway analysis.

To optimize docking scores, ligands are often forced into strained pocket-bound conformations, causing bond-order, charge, and valency misperception during hydrogen addition. Charge-delocalized ligands such as NADPH and ZMA docking poses often receive incorrect formal charges after bond-order correction, due to dense N/O/P heteroatom networks that are highly sensitive to small BO differences. In these cases, RDKit and Open Babel assign compensating charges, which must be corrected against the SMILES to obtain chemically accurate, simulation-ready structures. HAddiT corrected these inconsistencies, delivering structures ready for metabolism studies. The pictorial representation of correct hydrogen addition by HAddiT vs the conventional methods is shown below.

Future Directions

HAddiT will be equipped with the capability to assign protonation states using both conventional rule-based approaches and the more accurate Aganitha’s pKa estimation pipeline.