Research
Lumi: Conversational AI for Public Services
Lumi is a conversational AI prototype I designed to help international students find their way through Finnish public services. This case study covers what happened when I studied real newcomers using it, and what that revealed about how trust in AI actually works.
- Role
- Sole researcher and designer
- Team
- Supervised by Prof. Thomas Olsson and Amir Pakpour Haji Agha
- Duration
- 2025 – 2026
- Tools
- Figma (prototype), Atlas.ti
The problem
Moving to a new country means learning a system nobody explains to you. Where to register. Which office handles what. Which number on a form actually matters.
Finland's public services are, on paper, well organized. In practice, newcomers still have to piece the sequence together themselves: asking a friend, searching in English on a Finnish institution's website, or guessing.
One participant in my study put it simply. Talking about their first months in the country, they said it "often feels like you're just left here... you are kind of lost."
Often feels like you're just left here... you are kind of lost.
14 of the 15 people I studied described some version of the same thing: they wished something like Lumi had existed when they arrived.
The study
I designed Lumi to test a specific question: when an AI assistant helps someone through an unfamiliar, high-stakes system, what actually makes them trust it, or stop trusting it?
To find out, I recruited 15 international students at Tampere University who had personally navigated these systems as newcomers, and ran single-session studies with each of them. Every session lasted about 35 minutes and moved through four real scenarios: student housing, banking, healthcare, and university services. Before and after each session, I asked participants to describe what they felt, not just what they thought.
All sessions were consented and recorded, and every participant's data was anonymized before analysis. I analyzed the sessions using reflexive thematic analysis, a well-established qualitative method, and used an existing academic framework called Trust-as-Affect as a lens, not a theory I was trying to prove. Throughout, I looked for contradictions in the data as carefully as I looked for agreement, one wrong answer mattered as much as a pattern across many.
Two pilot sessions came first. They showed me exactly where the prototype broke, and that shaped everything I built next.
- 1
Pilot session 1
- 2
Pilot session 2
- 3
Main study
15 participants, single sessions, ~35 minutes each
Building Lumi
The first version of Lumi worked well when people clicked through its guided structure. It failed the moment someone typed a question outside that structure, and in the pilot sessions, that happened often enough to break the interaction entirely. If I'd run the main study on that version, I would have lost the exact moments I was there to study: real, unscripted moments of doubt and confidence, every time someone went off-script.
So I added a constrained AI layer underneath the structured design, one that could handle open-ended questions while keeping Lumi's tone and shape consistent. This carried its own risk: a second system, however well constrained, could easily have felt like two different assistants stitched together, or introduced new failures of its own. I tested it carefully before the main study began, and in practice, the transition held up. Typing a question outside the guided path felt like a natural continuation of the same conversation, not a handoff to something else.
Every screen followed the same design logic: real institutions named directly (TOAS for housing, Kela for benefits, YTHS for healthcare), links out to their official pages, and short, scannable summaries instead of long paragraphs. Then I watched what happened when real people relied on it.
What we found
Trust didn't behave the way I expected going in.
It formed slowly, in small moments: an official link that worked, a price range that matched what someone already expected, a summary that read as precise instead of vague. Each one added a little more confidence.
But trust could also disappear all at once. In one session, Lumi confidently named a housing provider that doesn't exist. The participant clicked the link it gave. Nothing loaded. That single moment was enough. Everything Lumi had earned up until then stopped mattering, and the participant's trust reset instantly, not gradually.
That asymmetry, slow to form, fast to break, turned out to be the clearest finding of the study.
A second pattern stood out almost as strongly. Vague reassurance didn't move people. Specific facts did. When Lumi gave one participant a real price range for student housing, they said it plainly: "I was immediately motivated... attracted to the price."
I was immediately motivated... attracted to the price.
Participants also stayed a little careful even when they trusted Lumi, checking its answers against another source out of habit, not suspicion, since they already knew AI systems can be wrong.
These patterns fell into three broad areas: what made trust possible in the first place, how people managed it moment to moment, and where it eventually broke down. Across the study, they organized into seven themes. (The full analysis behind these themes is reserved for future academic publication.)
What it means
The clearest output of this research is a set of seven design principles for building trustworthy conversational AI in high-stakes settings, not just for Lumi, but for any assistant operating in this kind of context. Each one traces back to something that actually happened in the study, not just intuition.
What made trust possible
What kept trust honest
Where trust needs protecting
This research also contributes to how trust itself is understood in human-AI interaction, extending an existing academic framework in ways I'm continuing to develop toward publication. The seven principles are what I'd build from that understanding today. The next test is building them into something people rely on before they ever have a reason to doubt it.
Reflection
This was a small study: 15 people, one university, one prototype. It doesn't describe how everyone experiences trust in AI, only how these 15 people did, in real depth.
I was also both the designer of Lumi and the researcher studying it, a position I stayed conscious of throughout, especially once the prototype's own mistakes became some of the most important data in the study.
If I ran this again, I'd recruit across more than one university, and stress-test the fallback layer harder before the real sessions began.
What stayed with me most is simple. An AI system doesn't earn trust once. It earns it constantly, and it can lose it instantly.
Go deeper
This project is part of my Master's thesis, completed at Tampere University and graded 5/5. I'm preparing parts of this study for academic publication, so a few details stay under wraps for now, but I'm glad to share more directly.
The full thesis and a working demo of the prototype are available on request.
Request the thesis & a demo