Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
From Language to Action: A System for Robotic Pick-and-Place Code Generation
University West, Department of Engineering Science.
2025 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Recent advancements in Artificial Intelligence, particularly in Large Language Models (LLMs) and Vision-Language Models (VLMs), have opened new possibilities for improving human-robot interaction. However, enabling robots to understand natural language instructions and accurately execute them in dynamic environments remains a major challenge. Issues like hallucinated output from language models and inaccurate object localization limit the use of AI in real-world robotic applications. This thesis addresses these challenges by proposing a system that generates action plans and robot control code for pick-and-place tasks based on natural language input.

The system integrates LLMs and VLMs into a three-layer framework: an Interpreter layer to analyze user queries, an Object Localizer layer to find the referred object in a simulated environment, and a Responder layer to generate an executable code. The system was tested using a UR5 robot in simulation, and two VLMs, CLIPSeg and Grounding DINO, were compared and modified for improved accuracy. Additionally, Retrieval- Augmented Generation (RAG) was implemented to provide external documents to the LLM, reducing hallucination, and improving code quality. The results show that the system can reliably interpret user queries and perform pick-and-place tasks, forming a solid foundation for expanding to more complex robotic behaviors.

Place, publisher, year, edition, pages
2025. , p. 50
Keywords [en]
Human-Robot-Interaction, Artificial Intelligence, Code Generator System, Large-Language Model, Vision-Language Model, Computer Vision
National Category
Robotics and automation
Identifiers
URN: urn:nbn:se:hv:diva-23921Local ID: EXC915OAI: oai:DiVA.org:hv-23921DiVA, id: diva2:1988346
Subject / course
Robotics
Educational program
Master i robotik och automation
Supervisors
Examiners
Available from: 2025-08-26 Created: 2025-08-11 Last updated: 2025-09-30Bibliographically approved

Open Access in DiVA

No full text in DiVA

By organisation
Department of Engineering Science
Robotics and automation

Search outside of DiVA

GoogleGoogle Scholar

urn-nbn

Altmetric score

urn-nbn
Total: 222 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf