Leonardo Solis Vasquez

Professor & Researcher

Leonardo
Solis-Vasquez

High-Performance Computing | Heterogeneous Computing | GPU & FPGA Acceleration

Universidad de Ingeniería y Tecnología (UTEC) · Lima, Peru

About Me

High-performance and heterogeneous computing for demanding scientific and engineering applications

I am a professor and researcher specializing in high-performance computing (HPC) and heterogeneous computing. I am currently with the Universidad de Ingeniería y Tecnología (UTEC) in Lima, Peru, where my work focuses on parallel computing, hardware acceleration, and efficient computing for scientific and engineering applications.

I hold a PhD in Computer Science from the Technical University of Darmstadt ↗, Germany. My doctoral research ↗ explored the acceleration of molecular docking through heterogeneous computing, including multicore CPUs, GPUs, and FPGAs, with particular attention to performance, accuracy, and energy efficiency.

My research and development experience spans GPU computing, FPGA acceleration, parallel programming, and performance optimization across heterogeneous architectures. I have worked with technologies such as CUDA, SYCL, and OpenCL and have participated in international research collaborations involving academic and industrial partners. My work has resulted in peer-reviewed publications in high-performance computing, heterogeneous systems, bioinformatics, and related areas.

I am an IEEE Senior Member and an active reviewer for international journals and conferences. I am particularly interested in projects involving computational performance, accelerator-based computing, and the efficient mapping of demanding applications to modern hardware.

Technical Expertise & Collaboration

High-Performance Computing & Performance Optimization

Performance analysis and optimization of compute-intensive applications on CPUs and GPUs, including profiling, parallelization, and efficient use of modern computing architectures.

Heterogeneous Computing & Software Portability

Development and migration of parallel applications across heterogeneous platforms using technologies such as CUDA, SYCL, and OpenCL, with a focus on performance and portability.

FPGA & Hardware Acceleration

Acceleration of computational workloads using FPGAs, high-level synthesis, and hardware/software co-design for scientific and engineering applications.

Research Projects

Apr. 2025 – Dec. 2025

Scale4Edge: as part of the "Future-Proof Special Processors and Development Platforms (ZuSE)" call for proposals of the BMBF's Digital Strategy.

Technical University of Darmstadt (Germany)

Scale4Edge, funded by the German Federal Ministry of Education and Research (BMBF), investigated methods and tools for the automated specialization of processors through application-specific instruction-set extensions. The project targeted improvements in computational performance and energy efficiency for domains such as artificial intelligence, automotive systems, and cryptography.

Jun. 2022 - Mar. 2025

Intel Center of Excellence: Porting AutoDock-GPU to oneAPI and Intel XPUs

Technical University of Darmstadt (Germany)

Funded by Intel through the oneAPI Center of Excellence at TU Darmstadt, this project investigated the portability of AutoDock-GPU across heterogeneous computing platforms. My work focused on porting the CUDA implementation to SYCL and evaluating and optimizing its performance on high-end Intel and NVIDIA GPUs, with the goal of maintaining a portable codebase across different accelerator architectures.

Mar. 2020 – May 2022

PANDAS: Programmable Appliance for Near-Data Processing Accelerated Storage

Technical University of Darmstadt (Germany)

Funded by the German Federal Ministry of Education and Research (BMBF), PANDAS investigated near-data processing architectures for database systems and computational storage. The project involved the design and optimization of hardware and software mechanisms for executing data-processing operations closer to storage, with particular attention to system architecture, performance, and database workloads.

Jul. 2020 – Mar. 2022

Accelerating Molecular Docking using Vector Computer Systems

Technical University of Darmstadt (Germany)

In collaboration with NEC Germany, this project investigated the porting and optimization of AutoDock molecular docking workloads for the NEC SX-Aurora TSUBASA Vector Engine. The work focused on adapting irregular computational kernels to vector architectures and evaluating their performance on this specialized high-performance computing platform.

Sep. 2018 – Feb. 2019

EPHoS: Evaluation of Programming Models for Heterogeneous Systems

Technical University of Darmstadt (Germany)

Funded by the German Association of the Automotive Industry (VDA), EPHoS investigated parallel programming models for heterogeneous embedded computing platforms in the automotive domain. Automotive and autonomous-driving workloads were used to evaluate performance, portability, and programmability across different execution platforms and parallel programming approaches.

Mar. 2016 – Sep. 2019

Hardware Acceleration of Molecular Docking (Ph.D. thesis)

Technical University of Darmstadt (Germany) and
The Scripps Research Institute (USA)

As part of my doctoral research, including a research stay at The Scripps Research Institute supported by DAAD and PRONABEC, I developed and evaluated parallel implementations of AutoDock molecular docking using OpenCL. The work investigated multicore CPUs, GPUs, and FPGAs and included a comprehensive evaluation of execution performance, quality of results, and computational energy efficiency across heterogeneous computing platforms.

Mar. 2014 – May 2015

A cost-effective and portable system for automatic diagnosis of pneumonia in children

Universidad Peruana Cayetano Heredia (Peru)

Supported by Grand Challenges Canada, this project investigated a portable system for assisting the diagnosis of pediatric pneumonia from lung ultrasound images. The work combined image-processing and pattern-recognition techniques with a portable acquisition platform intended for use by trained personnel in regions with limited access to medical specialists.

Mar. 2014 – May 2015

ePetri: Electronic Petri

Universidad Peruana Cayetano Heredia (Peru)

This project investigated a lens-free imaging system for assisting the interpretation of Microscopic Observation Drug Susceptibility (MODS) cultures for tuberculosis diagnosis. The system enabled microscopic imaging of biological samples over a relatively large field of view, with the goal of supporting portable and cost-effective diagnostic workflows.

Apr. 2013 – Oct. 2013

Parallelization of video streaming software

Politecnico di Torino (Italy)

This project investigated the use of an experimental software-analysis tool to identify data dependencies and performance bottlenecks in C applications. The resulting analysis was used to guide the parallelization of computational kernels for multicore CPU architectures.

Apr. 2012 – Nov. 2012

Set-Top Box Prototype providing Security Components on FPGA (M.Sc. thesis)

Center for Advanced Security Research Darmstadt (Germany)

As part of my Erasmus study exchange at TU Darmstadt, I modified an FPGA-based MPEG-2 decoder used in a set-top-box prototype to investigate hardware-level broadcast security mechanisms. The work examined the interaction between video-decoding components and cryptographic hardware within an FPGA-based system.

Nov. 2008 – Aug. 2010

Various electronic-engineering related projects (as B.Sc. graduate)

Universidad Nacional de Ingeniería (Peru)

During this period, I worked on several embedded and reconfigurable-computing projects, including the implementation of communication protocols on MSP430 microcontrollers, evaluation of voice-compression algorithms on Spartan-3 FPGAs, and development of fingerprint-identification algorithms on TMS320 DSPs. I also implemented a chaos-based Even-Mansour cipher on Cyclone IV FPGAs.

Publications

Peer-reviewed research papers in journals
Near-Data Processing in Database Systems on Native Computational Storage under HTAP Workloads
Sep. 2022

T. Vinçon, C. Knödler, L. Solis-Vasquez, A. Bernhardt, S. Tamimi, L. Weber, F. Stock, A. Koch, I. Petrov.

VLDB Endowment, Volume 15

Benchmarking the Performance of Irregular Computations in AutoDock-GPU Molecular Docking
Nov. 2021
L. Solis-Vasquez, A. F. Tillack, D. Santos-Martins, A. Koch, S. LeGrand, S. Forli.

Elsevier, Parallel Computing
Porting and Optimizing Molecular Docking onto the SX-Aurora TSUBASA Vector Computer
Sep. 2021
L. Solis-Vasquez, E. Focht, A. Koch.

South Ural State University, Supercomputing Frontiers and Innovations (JSFI)
On the necessity of explicit cross-layer data formats in near-data processing systems
Mar. 2021
L. Weber, T. Vinçon, C. Knödler, L. Solis-Vasquez, A. Bernhardt, I. Petrov, A. Koch. Springer, Distributed and Parallel Databases (DADP)
Accelerating AutoDock4 with GPUs and Gradient-Based Local Search
Jan. 2021
D. Santos-Martins, L. Solis-Vasquez, A. F. Tillack, M. F. Sanner, A. Koch, S. Forli.

ACS, Journal of Chemical Theory and Computation (JCTC)
nKV in Action: Accelerating KV-Stores on Native Computational Storage with Near-Data Processing
Aug. 2020
T. Vinçon, L. Weber, A. Bernhardt, C. Riegger, S. Hardock, C. Knoedler, F. Stock, L. Solis-Vasquez, S. Tamimi, A. Koch, I. Petrov.

VLDB Endowment, Volume 13
D3R Grand Challenge 4: prospective pose prediction of BACE1 ligands with AutoDock-GPU
Nov. 2019
D. Santos-Martins, J. Eberhardt, G. Bianco, L. Solis-Vasquez, F. A. Ambrosio, A. Koch, S. Forli.

Springer, Journal of Computer-Aided Molecular Design (JCAMD)
Comparison of affinity ranking using AutoDock-GPU and MM-GBSA scores for BACE-1 inhibitors in the D3R Grand Challenge 4
Nov. 2019
L. E. Khoury, D. Santos-Martins, S. Sasmal, J. Eberhardt, G. Bianco, F. A. Ambrosio, L. Solis-Vasquez, A. Koch, S. Forli, D. L. Mobley.

Springer, Journal of Computer-Aided Molecular Design (JCAMD)
Automatic classification of pediatric pneumonia based on lung ultrasound pattern recognition
Dec. 2018
M. Correa, M. Zimic, F. Barrientos, R. Barrientos, A. Román-Gonzalez, M. J. Pajuelo, C. Anticona, H. Mayta, A. Alva, L. Solis-Vasquez, D. A. Figueroa, M. A. Chavez, R. Lavarello, B. Castañeda, V. A. Paz-Soldán, W. Checkley, R. H. Gilman, R. Oberhelman.

Public Library of Science, PLOS ONE Journal
Evaluation of a lens-free imager to facilitate tuberculosis diagnostics in MODS
Dec. 2015
L. Solis, J. Coronel, D. Rueda, R. H. Gilman, P. Sheen, M. Zimic.

Elsevier, Journal of Tuberculosis
Peer-reviewed research papers in conferences
Introduction to Parallel Computing Using CPUs and GPUs
Mar. 2026

L. Solis-Vasquez.

IEEE, 14th International Conference on Software Process Improvement (CIMPS)

GPU-Accelerated Drug Discovery with Docking on the Summit Supercomputer: Porting, Optimization, and Application to COVID-19 Research
Jul. 2020

S. LeGrand, A. Scheinberg, A. F. Tillack, M. Thavappiragasam, J. V. Vermaas, R. Agarwal, J. Larkin, D. Poole, D. SantosMartins, L. Solis-Vasquez, A. Koch, S. Forli, O. Hernandez, J. C. Smith, A. Sedova.

ACM, 11th International Conference on Bioinformatics, Computational Biology and Health Informatics (BCB)

Evaluating the Energy Efficiency of OpenCL-accelerated AutoDock Molecular Docking
Mar. 2020
L. Solis-Vasquez, D. Santos-Martins, A. Koch, S. Forli.

IEEE, 28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)
Using Parallel Programming Models for Automotive Workloads on Heterogeneous Systems - a Case Study
Mar. 2020

L. Sommer, F. Stock, L. Solis-Vasquez, A. Koch.


IEEE, 28th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)

Work-in-Progress: DAPHNE - An Automotive Benchmark Suite for Parallel Programming Models on Embedded Heterogeneous Platforms
Oct. 2019
L. Sommer, F. Stock, L. Solis-Vasquez, A. Koch.

ACM, International Conference on Embedded Software Companion (EMSOFT)
Filtering of the skin portion on lung ultrasound digital images to facilitate automatic diagnostics of pneumonia
Nov. 2016
F. Barrientos, A. Román-Gonzalez, R. Barrientos, L. Solis, A. Alva, M. Correa, M. Pajuelo, C. Anticona, R. Lavarello, B. Castañeda, R. Oberhelman, R. H. Gilman, M. Zimic.

IEEE, XXXVI Convención de Centro América y Panamá (CONCAPAN)
Automatic detection of pneumonia analyzing ultrasound digital images
Nov. 2016
R. Barrientos, A. Román-Gonzalez, F. Barrientos, L. Solis, M. Correa, M. Pajuelo, C. Anticona, R. Lavarello, B. Castañeda, R. Oberhelman, W. Checkley, R. H. Gilman, M. Zimic.

IEEE, XXXVI Convención de Centro América y Panamá (CONCAPAN)
Peer-reviewed research papers in symposiums and workshops
A Compute Graph Simulation and Implementation Framework Targeting AMD Versal AI Engines
Nov. 2025

J. Strobl, L. Solis-Vasquez, Y. Lavan, A. Koch.


ACM, SC ’25 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis

Architecting Tensor Core-Based Reductions for Irregular Molecular Docking Kernels
Nov. 2025

L. Solis-Vasquez, A. F. Tillack, D. Santos-Martins, A. Koch, S. Forli.

ACM, SC ’25 Workshops of The International Conference for High Performance Computing, Networking, Storage and Analysis.

Graphtoy: Fast Software Simulation of Applications for AMD’s AI Engines
Mar. 2024

J. Strobl, L. Solis-Vasquez, Y. Lavan, A. Koch.


Springer, 20th International Symposium on Applied Reconfigurable Computing (ARC)

Altis-SYCL: Migrating Altis Benchmarking Suite from CUDA to SYCL for GPUs and FPGAs
Nov. 2023
C. Weckert, L. Solis-Vasquez, J. Oppermann, A. Koch, O. Sinnen.

ACM, SC ’23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis
Experiences Migrating CUDA to SYCL: A Molecular Docking Case Study
Apr. 2023
L. Solis-Vasquez, E. Mascarenhas, A. Koch.

ACM, 11th International Workshop on OpenCL (IWOCL)
Simulating Molecular Docking on the SX-Aurora TSUBASA Vector Engine
Feb. 2023
L. Solis-Vasquez, E. Focht, A. Koch.

Springer, Joint Workshop on Sustained Simulation Performance (WSSP 2021)
Mapping Irregular Computations for Molecular Docking to the SX-Aurora TSUBASA Vector Engine
Nov. 2021
L. Solis-Vasquez, E. Focht, A. Koch.

IEEE, 11th Workshop on Irregular Applications: Architectures and Algorithms (IA3)
Result-Set Management for NDP Operations on Smart Storage
Jun. 2022
T. Vinçon, C. Knödler, A. Bernhardt, L. Solis-Vasquez, L. Weber, A. Koch, I. Petrov.

ACM, 18th International Workshop on Data Management on New Hardware (DaMoN)
A Framework for the Automatic Generation of FPGA-based Near-Data Processing Accelerators in Smart Storage Systems
May 2021
L. Weber, L. Sommer, L. Solis-Vasquez, T. Vinçon, C. Knödler, A. Bernhardt, I. Petrov, A. Koch.

IEEE, International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
A cost model for NDP-aware query optimization for KV-stores
Jun. 2021
C. Knödler, T. Vinçon, A. Bernhardt, L. Solis-Vasquez, L. Weber, A. Koch, I. Petrov.

ACM, 17th International Workshop on Data Management on New Hardware (DaMoN)
Parallelizing Irregular Computations for Molecular Docking
Nov. 2020
L. Solis-Vasquez, D. Santos-Martins, A. F. Tillack, A. Koch, J. Eberhardt, S. Forli.

IEEE, 1010th Workshop on Irregular Applications: Architectures and Algorithms (IA3)
A Case Study in Using OpenCL on FPGAs: Creating an Open-Source Accelerator of the AutoDock Molecular Docking Software
Aug. 2018
L. Solis-Vasquez, A. Koch.

IEEE, 5th International Workshop on FPGAs for Software Programmers (FSP)
A Performance and Energy Evaluation of OpenCL-accelerated Molecular Docking
May 2017
L. Solis-Vasquez, A. Koch.

ACM, 5th International Workshop on OpenCL (IWOCL)
Miscellaneous research articles and contributions
Speeding up simulations for drug discovery with Autodock-GPU
Oct. 2021
L. Solis-Vasquez, A. F. Tillack, D. Santos-Martins, A. Koch, S. Forli.

HiPEAC, HiPEAC info 64
DAPHNE - An Automotive Benchmark Suite for Parallel Programming Models on Embedded Heterogeneous Platforms
Apr. 2020
L. Sommer, F. Stock, L. Solis-Vasquez, A. Koch.

Future Automotive HW/SW Platform Design (Dagstuhl Seminar 19502), Dagstuhl Reports
EPHoS: Evaluation of Programming Models for Heterogeneous Systems
Jun. 2019
L. Sommer, F. Stock, L. Solis-Vasquez, A. Koch.

German Association of the Automotive Industry (VDA)
Test set of 140 complexes for AutoDock-GPU [Data set]
Sep. 2020
D. Santos-Martins, L. Solis-Vasquez, A. F. Tillack, M. F. Sanner, A. Koch, S. Forli.

Zenodo
Theses
Accelerating Molecular Docking by Parallelized Heterogeneous Computing - A Case Study of Performance, Quality of Results, and Energy-Efficiency using CPUs, GPUs, and FPGAs
Dec. 2019
L. Solis Vasquez.

TUprints, Ph.D. Thesis
Implementation of a Set-Top Box Architecture providing Security Components on an FPGA
Mar. 2013
L. Solis Vasquez.

Politecnico di Torino, M.Sc. Thesis

Get In Touch

Whether you’d like to discuss a research collaboration, technical project, or new idea, feel free to get in touch using the form below.