back to list

Project: [ATOS] Generative AI projects

Description
ATOS offers a number of interesting projects:

Evaluating Multi-Agent Software Engineering Systems in Enterprise Development Environments

Problem Definition
Recent advances in Large Language Models (LLMs) and AI agents have enabled systems that can perform software engineering tasks such as requirements analysis, code generation, testing, debugging, and code reviews. While these systems show promising results on isolated tasks, their effectiveness in end-to-end software development workflows remains poorly understood.

Organizations are interested in understanding how far autonomous software engineering can be taken in practice, and which software development activities can be safely delegated to AI agents while maintaining quality, security, and compliance standards. It remains unclear which level of autonomy can realistically be achieved and how autonomous software engineering agents should be evaluated.

Objective
Investigate the feasibility, reliability, and limitations of autonomous software engineering agents in realistic software development environments.

The student will design and implement an experimental framework in which AI agents perform software engineering tasks with varying levels of autonomy. The framework will be used to evaluate the performance of agents across different development scenarios and investigate factors that contribute to success and failure.

Potential topics include:
· Autonomous feature implementation
· Automated bug fixing
· Agent-driven software testing
· Multi-agent collaboration during software development
· Human-AI collaboration models
· Self-improving software engineering agents

Research Questions
· Which software engineering activities can be reliably performed by autonomous AI agents?
· How does agent autonomy impact software quality, productivity, and maintainability?
· What are the most common failure modes of autonomous software engineering agents?
· How can autonomous software engineering systems be evaluated beyond traditional coding benchmarks?
· How do single-agent and multi-agent approaches compare for complex software engineering tasks?
· What level of human oversight is required to ensure trustworthy operation?

Scientific Contribution
Possible scientific contributions include:
· A novel evaluation framework for autonomous software engineering systems
· New metrics for measuring agent autonomy and reliability
· Empirical comparison of single-agent and multi-agent software engineering approaches
· Identification and categorization of failure modes in autonomous development workflows
· Guidelines for designing trustworthy software engineering agents in enterprise environments

Deliverables
· Literature review on AI agents and autonomous software engineering
· Experimental framework and benchmark scenarios
· Prototype implementation of one or more autonomous software engineering agents
· Empirical evaluation using quantitative and qualitative metrics
· Analysis of limitations, risks, and future research directions
· Final thesis report and presentation


Designing a Way of Working for AI Token Monitoring and Process Optimization

Problem Definition
Generative AI systems are increasingly used in software development and business processes, creating new cost-management challenges. Unlike traditional cloud costs, token usage is often variable, difficult to attribute, and influenced by prompts, context size, model selection, output length, and retry behavior.

Organizations need a way to monitor and optimize AI usage without slowing down experimentation, adoption, or developer productivity. This thesis investigates how an AI FinOps way of working can create cost transparency and responsible usage while preserving business value.

Objective
Develop a practical operating model and measurement framework for AI token monitoring, cost attribution, and process optimization in enterprise environments.

The student will analyze AI usage patterns, identify token waste drivers, and design governance mechanisms such as showback dashboards, model tiering, guardrails, and feedback loops that support responsible AI adoption.

Potential topics include:
· Token monitoring and cost attribution
· Prompt and context optimization
· Model tiering and usage policies
· Cost-per-outcome metrics
· Governance models for AI adoption

Research Questions
· How can token usage be monitored and attributed in enterprise AI environments?
· Which usage patterns create the largest token waste?
· How can organizations reduce AI costs without limiting productivity or adoption?
· Which governance mechanisms support responsible and effective AI usage?

Scientific Contribution
· A framework for AI FinOps in token-based and agentic AI workloads
· A taxonomy of token waste drivers in enterprise AI usage
· A measurement model linking token cost to productivity and business value
· Design guidelines for governance mechanisms that improve transparency without blocking useful AI adoption

Deliverables
· Literature review on AI FinOps and GenAI cost management
· AI FinOps operating model and measurement framework
· Analysis of token waste drivers and optimization opportunities
· Empirical evaluation using selected AI workflows or proof-of-concept scenarios
· Final thesis report and presentation


Responsible-AI Gate for CI/CD: EU-AI-Act Ready MLOps Pipeline

Problem Definition
Organizations are scaling GenAI and ML, but often lack governance, compliance, and auditability in their MLOps pipelines. The EU AI Act increases the need for demonstrable control mechanisms, while standard pipelines typically do not include automated quality and risk checks.

Objective
Design and implement a Responsible-AI Gate that automatically evaluates data quality, bias, fairness, model behavior, explainability, documentation, audit trails, and deployment approval criteria in a CI/CD pipeline.

Potential topics include:
· Responsible-AI metrics
· MLOps governance
· Bias, fairness, and safety checks
· Auditability and deployment approval

Research Questions
· Which Responsible-AI metrics are minimally required based on NIST and the EU AI Act?
· How can Responsible-AI checks be integrated into DevOps without reducing delivery speed?
· What are the trade-offs between reliability, cost, and latency?

Deliverables
· Responsible-AI Gate module
· Python evaluation scripts for bias, fairness, and safety
· Governance dashboard
· Model card template and architecture document
· Final thesis report with impact analysis

Unified Evaluation Harness for RAG, LLM Agents & Copilot Scenarios

Problem Definition
GenAI systems can show inconsistent quality due to hallucinations, weak retrieval, variable latency, and changing costs. Organizations need predictable KPIs and a reliable evaluation framework for RAG, LLM agents, and Copilot scenarios.

Objective
Develop an evaluation harness that automatically tests GenAI systems on groundedness, correctness, hallucination detection, retrieval quality, latency, token costs, robustness, and regression behavior after model updates.

Potential topics include:
· Groundedness and hallucination metrics
· RAG evaluation
· Regression testing for LLMs
· Latency and token-cost analysis

Research Questions
· How can reliable groundedness and hallucination metrics be defined?
· How can RAG systems be evaluated systematically?
· Which techniques are suitable for LLM regression testing?
· Where are the main opportunities for cost optimization?

Deliverables
· Python evaluation framework
· Benchmark dataset and judge prompts
· Test report generator for quality, risk, and cost
· CI integration for regression tests
· Final thesis report with recommendations

LLM-Driven Development Cycle Manager

Problem Definition
Software teams use AI tools such as GitHub Copilot and Claude Code, but often lack an integrated workflow that connects code repositories, ticketing systems, bug flows, and pull request generation.

Objective
Develop a prototype of an LLM-driven development workflow manager that reacts to GitHub events, integrates with Jira or Azure DevOps, links bugs to code locations, and generates initial pull request proposals using an LLM.

Potential topics include:
· LLM-based code patch generation
· GitHub event automation
· Integration with Jira or Azure DevOps
· Safety checks for automated pull requests

Research Questions
· How reliable are LLMs in generating small, controlled code patches for bug fixes?
· Which GitHub event triggers provide the most value for automation?
· How can unwanted changes in automated pull requests be prevented?
· What are the relevant limitations regarding security, tokens, permissions, and rate limits?

Deliverables
· Prototype LLM Development Cycle Manager
· Observability dashboard
· Cost governance playbook
· Final thesis report and demonstration environment
Details
Supervisor
Joaquin Vanschoren
External location
ATOS
Interested?
Get in contact