Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Creating a GPU Cluster for AI Model Training using Rancher Kubernetes and KAI Scheduler
University West, Department of Engineering Science.
University West, Department of Engineering Science.
2025 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

This thesis describes the implementation of a GPU-cluster using Rancher Kubernetes and KAI Scheduler. After conducting thorough research and preparations related to require-ments, tools and circumstances, a functional cluster was implemented as a proof of con-cept. The cluster is designed to enable students and researchers at University West to train AI models using Machine Learning.

Implementation included configuring Kubernetes to optimize GPU utilization and integrat-ing KAI Scheduler to enable several jobs to simultaneously run on the same GPU. NVIDIA GPU Operator was also used to automate the installation of necessary software components, simplifying GPU availability and enabling easier scalability when adding more nodes to the cluster.

To align with University West’s vision of the cluster, several important features were imple-mented, including checkpoints for saving progress in case of unexpected interruptions, and job splitting to run segments in parallel instead of in sequence, potentially reducing training times for large jobs. While these features are implemented and tested, their usability could be improved in future work.

The findings of this thesis highlight the cluster's potential for enhancing AI model training at University West. Although the cluster is not usable at scale, the objective of developing a proof of concept was achieved as all specified requirements were met, laying a solid foun-dation for future work to improve the cluster further. Date:

Place, publisher, year, edition, pages
2025. , p. 51
Keywords [en]
GPU-cluster, AI-model training, Kubernetes, Rancher, KAI Scheduler, Machine Learning.
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:hv:diva-24050Local ID: EHD500OAI: oai:DiVA.org:hv-24050DiVA, id: diva2:1992643
Subject / course
Computer engineering
Educational program
Datateknik - högskoleingenjör
Supervisors
Examiners
Available from: 2025-09-03 Created: 2025-08-28 Last updated: 2025-09-30Bibliographically approved

Open Access in DiVA

No full text in DiVA

By organisation
Department of Engineering Science
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar

urn-nbn

Altmetric score

urn-nbn
Total: 82 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf