Roswell Park is seeking a partner to provide a high performance computing solution. We are not seeking a broker or consultant, unless that is the only method in which the solution provider does business. If you are interested in receiving the complete RFP package, please email Rob Henschel @
[email protected]
Roswell Park Comprehensive Cancer Center operates a shared research computing environment serving multiple academic research groups across translational science, bioinformatics, biostatistics, genomics, and machine learning disciplines. The platform supports several independent research groups, typically consisting of approximately five to eight users per group, with bioinformatics teams representing the heaviest computational users.
Current annual compute utilization is approximately 100,000 CPU core hours. Workloads are primarily CPU-based, with growing interest in GPU-accelerated machine learning and AI workflows. Most jobs are Slurm-scheduled and consist of R, Python, and genomics pipelines. There is increasing interest from research groups in higher-level, platform-based approaches that enable execution of complex genomics workflows (e.g., sequencing processing, variant analysis) without requiring direct interaction with underlying infrastructure. Workloads are largely embarrassingly parallel in nature, and job durations commonly range from several hours to multiple days.
Queueing is acceptable and expected. Roswell Park prefers high resource utilization with reasonable queue wait times over over-provisioned idle infrastructure. Jobs running for 24 hours or more are common, and short wait times of several hours are operationally acceptable. However, the environment must ensure that high-demand workloads can complete within reasonable timeframes under typical peak usage.
The previous solution provided access to approximately 21 CPU nodes and approximately 5 GPU-enabled nodes, including NVIDIA A100 and legacy V100 capacity. Respondents should map proposed configurations to this baseline and explain any material differences in concurrency, memory per core, queueing behavior, GPU availability, and storage performance.
Storage requirements currently approach one petabyte of high throughput (non-cold). Respondents should map their proposed baseline configuration to this requirement. Respondents may demonstrate, in their cost optimized configuration, thoughtful storage tiering strategies to balance performance and cost efficiency, particularly given that a portion of stored data may not require immediate high-speed access.
Roswell Park’s current general research HPC environment is not assumed to host PHI as part of ordinary baseline operation. However, Roswell Park anticipates that future research computing needs may include PHI, NYS-regulated data, NIH controlled-access data, regulated genomic data, or other sensitive datasets requiring PHI-equivalent or enhanced security handling.
Roswell Park values predictable, budgeted cost models. The institution seeks a partner capable of recommending right-sized architectures that saturate compute effectively while preserving flexibility for future growth in GPU demand, storage expansion, and workload diversification.
Business enterprises awarded an identical or substantially similar procurement contract within the past five years:
University at Buffalo Center for Computational Research