Dr. David E. Konerding, Ph.D.
Scientific ML Infrastructure | Data Engineering | Scalable Research Platforms
π§ dek@konerding.com | π San Francisco Bay Area, CA
Executive Summary
Engineering leader and hands-on engineer specializing in ultra-scalable infrastructure for scientific machine learning and high-performance computing. Combines deep domain expertise in biophysics and computational biology with decades of systems engineering at Google, Genentech, Insitro, and LBNL. Proven track record of managing multi-million-dollar compute budgets, leading engineering teams, and building platforms that accelerate research for hundreds of scientists.
Education
-
University of California, Berkeley | Postdoctoral Scholar (Brenner and Sjolander labs) | 2000 β 2003
-
University of California, San Francisco | Ph.D., Biophysics | 1995 β 2001
-
University of California, Santa Cruz | B.A., Biochemistry and Molecular Biology | 1991 β 1995
Core Competencies
- Scientific ML Infrastructure & MLOps
- Scalable Research & Compute Platforms
- High-Performance Computing (HPC) & Batch Queuing
- Hybrid Cloud & On-Prem Architecture
- Large-Scale TPU/GPU Fleet Health & Optimization
- Technical Leadership & Team Management
Technical Skills
Languages: Python, C++, SQL, Java, Bash
Compute & ML Infrastructure: Distributed Training Systems, TPU/GPU Accelerator Health, AWS/GCP, Batch Queuing Systems (Slurm, Grid Engine), Cluster Design & Performance Tuning
Data & Systems: Relational Databases (MySQL, PostgreSQL), High-Performance Networked File Systems, Hybrid Cloud Migration, Linux Kernel & OS Customization
Professional Experience
Genentech
Director, Machine Learning Engineer, gRED AI4DD
Jan 2026 β Present
- Manage a team of 3 software engineers developing core Machine Learning Infrastructure for the AI for Drug Discovery (AI4DD) organization.
- Oversee architecture and operational strategy for one of the largest computational allocations in the pharmaceutical industryβserving the needs of several hundred research scientists.
Senior Principal Machine Learning Engineer, AI4DD
Dec 2024 β Jan 2026
- Served as individual contributor tech lead engineering high-throughput machine learning infrastructure supporting several hundred researchers across AI4DD.
Principal Data Engineer, gRED Infrastructure & Architecture
Jun 2021 β Dec 2024
- Spearheaded gRED's enterprise on-premises to cloud infrastructure migration strategy, expanding compute capacity for scientific workloads.
- Diagnosed and resolved critical bottlenecks in complex research code bases to drastically improve execution speed and resource efficiency.
- Enhanced institutional and corporate support frameworks for large-scale research computing environments.
Senior System Architect for Research Computing
Aug 2007 β Aug 2008
- Served as the primary technical liaison bridging Genentech's Research and IT divisions, aligning infrastructure capabilities with domain-specific computational biology demands.
Google Inc.
Staff Engineer, Platforms Machine Learning
Aug 2019 β Jun 2021
- Led hardware health and reliability initiatives across Google's TPU fleet, engineering automated eviction and fault-detection protocols for unhealthy hardware to ensure uninterrupted job execution for research teams.
Staff Engineer, Infrastructure
Dec 2013 β Aug 2018
- Co-created and launched Google Cloud Genomics, architecting Google Cloud's foundational platform for processing and analyzing planetary-scale genomic datasets.
- Contributed to Sibyl, a machine learning product used to maximize revenue for Youtube, Play Store, and other Google products
Senior Engineer, Infrastructure
Aug 2009 β Dec 2013
- Architected and deployed a massive distributed computing system (Google Exacycle) that harvested idle compute cycles across global data centers for large-scale scientific simulations without impacting production serving.
Senior Test Engineer, Ads Database
Aug 2008 β Aug 2009
- Managed operations, scaling, and performance reliability for a high-throughput, revenue-critical core database.
Insitro
Senior Lead Data Engineer & Interim Head of Data Engineering
Aug 2018 β Aug 2019
- Designed and deployed Insitro's initial on-premises and cloud HPC and data engineering platforms from the ground up to empower early ML-driven drug discovery pipelines.
- Built and managed the foundational Data Engineering team, establishing engineering standards and driving executive recruitment.
Lawrence Berkeley National Laboratory
Computer Scientist, Computational Research Division
Aug 2003 β Aug 2007
- Developed pyGlobus, a widely used scientific grid computing framework enabling distributed data and compute management.
- Conducted independent research on applying distributed systems principles to complex high-performance scientific computing problems.
Publications
-
Ramsundar B, Kearnes S, Riley P, Webster D, Konerding DE, Pande V. Massively multitask networks for drug discovery. arXiv preprint arXiv:1502.02072. 2015.
-
Kohlhoff KJ, Shukla D, Lawrenz M, Bowman GR, Konerding DE, Belov D, Altman R, Pande VS. Cloud-based simulations on Google Exacycle reveal ligand modulation of GPCR activation pathways. Nat Chem. 2014;6(1):15-21.
-
Conway P, Tyka MD, DiMaio F, Konerding DE, Baker D. Relaxation of backbone bond geometry improves protein energy landscape modeling. Protein Sci. 2014;23(1):47-55.
-
Konerding DE, Cheatham TE 3rd, Kollman PA, James TL. Restrained molecular dynamics of solvated duplex DNA using the particle mesh Ewald method. J Biomol NMR. 1999;13(2):119-131.
References
Available upon request