I build and run software by directing AI coding agents. I write the rules and the spec, split the work, demand proof that it works, and operate what ships. When something breaks, I measure before I fix.
I don’t write most of the code by hand, and I won’t pretend to. What I own is whether it works.
Now
16, student at HTL Klagenfurt Mössingerstraße, electronics and computer engineering, year 3 of a 5-year technical college (graduating 2029)
Looking for
A summer internship in 2027, from 12 July, 4 to 8 weeks, in Klagenfurt or within a daily commute
Based in
Techelsberg am Wörthersee, Carinthia, Austria
Languages
German (native), English (fluent), Spanish (basic)
My own product. Windows automation for AI agents, on a desktop you can’t see. Private beta, started August 2026, launch planned for early 2027. agentfenster.com
AI agents that operate desktop apps usually take over your screen and mouse, or run in a cloud VM that doesn’t have your logins. Worse, they report success when a click hit nothing. agentfenster gives the agent its own hidden Windows desktop inside your session: same apps, same logins, and you can watch it live or take over. Claude Code and Codex drive it through an MCP server.
The rule the whole product is built around: never a silent success. After every action it reads the window again. If nothing changed, it says so and stops instead of clicking on.
AgentClaude Code or Codex asks for an action, for example “click Save”.
MCP serverFinds the element in the window’s accessibility tree, not in a screenshot.
Hidden desktopSends the action to the window itself. Your mouse and keyboard are never used.
CheckReads the window again and compares. Did the element react, did the window change?
ReportChanged: names the element it hit. Unchanged: fails loudly and stops.
One action, end to end. The check step is what makes it observable instead of hopeful.
Measured on my laptop while hardening the beta (August to October 2026)
What
Before
After
Time per click
2.5–7.7 s
0.1–0.5 s
Opening Edge on the hidden desktop
4.9 s
0.8 s
CPU with three people watching the live view
125 %
29 %
MCP server CPU while idle
132 %
0.6 %
Tokens the agent pays per session before doing anything
3,700
1,900
CPU is per-core percent, so 125 % means one and a quarter cores.
What I did
Came up with the product and its rules, for example “agents must never take over my mouse”, written down as hard project rules every agent reads first.
Decided the architecture questions with the agents, including what Windows can’t do (store apps open on the user’s desktop, the clipboard is shared) and stating those limits openly.
Used it for my own GUI work every day and logged every failure: about 95 detailed bug reports and 163 run reports so far.
Set the bar: a real task has to work end to end, not just a green test suite.
What the agents did
Wrote the code: Python, Win32 and UI Automation, a WebView2 interface, the MCP server.
Wrote the tests: over 9,500 automated tests (collected on 8 October 2026).
Measured with profilers and timers before changing anything, then proved the fix with the same measurement.
2,389 commits between 6 August and 7 October 2026.
Where it went wrong
On 5 October agentfenster reported that a text field “holds the text” while the field was actually empty: exactly the silent success the product exists to prevent. It was caught during daily use, went into the bug log, and the fix landed in the same session with a test that reproduces it. A green suite had not caught it, which is why real tasks are part of the definition of done.
The code is in a private repository because agentfenster is meant to become a paid product. I’m happy to walk through it, the architecture notes and the bug log in a call.
Running a production web app
Operations for a small coaching company (our family business). Live since 15 September 2026, several hundred customers.
The company sells a members’ web app: access codes, checkout, generated PDFs, a journal. I’m responsible for keeping it running: deploys, customer problems, and a support service I set up next to it. The support service sorts incoming requests with one model, drafts replies with a stronger one, and reproduces reported bugs in a separate copy of the app before anything touches production. It has 386 automated tests and an emergency stop: one file on the server switches it off.
Where the AI was wrong
A customer’s app kept crashing on her iPhone. The bug-fixing agent reported “fixed” four times. The usage data said otherwise: she kept bouncing back into the app within 40 seconds of opening one section. So we stopped trusting the agent’s claim and added diagnostic events around that exact screen. As the first fix, the heavy world map on that page now loads on iPhone only when tapped. The lesson: “fixed” needs evidence from real usage, not the agent’s word.
Stack: Node in Docker, SQLite, Caddy, a Python service under systemd, push alerts, automatic failover between AI accounts. I don’t show screenshots or names here because it’s a customer-facing system.
Mitschrift, an AI note-taker for my class
Side project at school. In daily use by two classmates and me since late September 2026.
Every lesson becomes one living note built from photos of the board, typed text, handwriting and, where allowed, audio. Audio is transcribed live in one-minute chunks, a short summary updates every minute, and a final note is written at the end of the lesson. At night my laptop re-transcribes the audio with a local model on the GPU for better accuracy.
Two decisions I made before writing a line of spec: audio is only recorded with the teacher’s permission, which is logged, and raw audio is deleted after seven days.
A problem worth knowing about: speech recognition models hallucinate on classroom noise. Ours kept producing the sentence “Untertitelung des ZDF” (subtitles of a German TV channel) when nobody said it, because it learned that phrase from TV subtitles. Transcripts are now checked against a list of these known phrases, and English filler from YouTube outros, before anything reaches the note.
Stack: FastAPI and SQLite, a Svelte web app that installs like an app, an Android app with its own recording service. 57 commits since 22 September 2026. Planned entry for Jugend Innovativ 2026/27.
Measured, then fixed
Four real problems from my own machines, written up like incident traces. In each one the agents measured first and changed things second.
Symptom
Measurement
Fix
Fix that didn’t work
Result
The laptop froze for seconds at a time
9 to 15 August 2026
Aug 9Aug 11Aug 13Aug 15
Symptom
Mouse and keyboard freeze for seconds, video editing looks like it crashes. Task Manager shows nothing unusual: no process is using the time.
Measure
The Windows hardware error log (WHEA) has 324 corrected PCIe errors in 14 days, all from the same PCIe port. The time is going into interrupt handling, which Task Manager doesn’t attribute to any process.
Fix
Replace the 2017 driver of the Realtek network chip on that port and switch off all its power saving. Next day: 845 errors in 30 seconds. Didn’t work.
Fix
Switch off PCIe link power management (ASPM). Errors continue, all “receiver errors”: the link is electrically marginal and the OS can’t heal it.
Fix
Hand PCIe error handling to the firmware with bcdedit /set {current} pciexpress forcedisable, so the OS no longer drowns in error interrupts.
Result
0 hardware errors over the next four days. A “crash” reported on day six turned out to have no crash artifact at all: it was Windows Update and the search indexer catching up.
What I took from it: before treating something as a crash, check whether a crash artifact exists.
Crashes that looked like a memory leak
4 to 7 August 2026
Aug 4Aug 5Aug 6Aug 7
Symptom
The laptop crashes eight minutes after boot. The suspicion is a memory leak.
Measure
No leak. A tool restored 13 agent sessions at once, 16 in total. Each started about 70 helper processes (MCP servers) and committed about 5 GB: 81 of 101 GB committed memory.
Fix
Cut the default helper servers from 18 to 8, the full set loads on demand. 5.1 to 3.24 GB per session, measured. Added a watchdog that cleans up orphaned processes.
Symptom
The next evening the watchdog fires at 74 % memory. It can’t fix it.
Measure
Ten restored sessions had not written a single byte in 12 hours but still held about 800 MB each. Idle sessions aren’t orphans, so the watchdog ignored them. A note-taking helper also launched up to 15 one-question AI calls per run, each starting the full set of 19 helper processes.
Fix
The watchdog now ends sessions idle for 90 minutes together with their whole process tree, and one-question calls start with no helpers at all: 1 extra process instead of 19.
Result
Committed memory 75 to 53 GB that evening. Two days later the same pattern: 195 idle processes removed, 95 % to 80 %.
What I took from it: “out of memory” was a fan-out problem, not a leak. The fix that mattered was in the rule for what counts as idle, not in the amount of RAM.
Python at 3,000 % CPU
17 September 2026, about one hour
Symptom
During a text-recognition batch job the laptop is close to unresponsive. Python processes use about 3,000 % CPU.
Measure
Ten worker processes at about 300 % each. The OCR library runs on onnxruntime, which by default starts a thread per core in every worker and ignores the usual OMP_NUM_THREADS setting. Ten workers times every core is oversubscription.
Fix
One thread per worker, at most six workers, lower process priority. Measured detail: child processes did not inherit the lower priority (parent “below normal”, workers “normal”), so it’s now set inside each worker.
Result
The job runs in the background while the laptop stays usable. The rule now applies to every parallel ML job: set threads per worker explicitly.
Console windows popping up on my screen
6 September 2026
Symptom
For about ten minutes, terminal windows keep popping up on my screen while agents work in the background.
Measure
A window watcher plus a controlled test: start a plain console directly on agentfenster’s hidden desktop and count what appears on mine. 2 windows. Windows 11 hands every new console to Windows Terminal, whose single process lives on my desktop.
Fix
Default terminal back to the classic console host, plus one launcher that starts scheduled tasks without a window. Every change reversible, with backups.
Result
Same test, 0 windows.
Bars show the order of the work. Where there is a date axis, positions are approximate within a day.
How I work
My projects are written by AI coding agents, mostly Claude Code, often several in parallel. That only works with structure. This is the loop I run.
Write the rules before the code. Each project has a rules file that every agent reads first: what the product promises, what must never happen, and the limits we have already measured. agentfenster’s starts with seven rules, the first one being “never a silent success”.
Split the work. Parallel agents get separate work packages and separate copies of the repository, so they don’t overwrite each other. I decide the order and what waits.
Demand evidence. A change isn’t done because the agent says so. It needs tests, a measurement, or a real run, and a skipped test counts as a finding, not as green.
Use it myself. I run my own tools daily. Every failure goes into a log with the same fields: expected, happened, workaround, fix, status.
Operate it. Linux servers, Docker, systemd, cron jobs, Cloudflare, push alerts to my phone when something needs me. When it breaks, back to step one with what we learned.
As long as that doesn’t work yet, it isn’t a mature product.
What I can do
Plan software and break it into work an agent can verify
Review changes, read logs, measure and debug on Windows and Linux
Run servers and deployments (Linux, Docker, Caddy, Cloudflare)
Electronics from school: circuits, lab measurements, Multisim
Video: filming, drone, editing (Premiere Pro)
What I can’t do yet
Write Python or TypeScript by hand. I’m learning to read the code I ship.
Java at school level, my third year of it.
Large-scale systems with real traffic. That’s exactly what I’d like to learn in an internship.
Also built
Video editing pipelinePersonal, in use
Agents cut my mountain-bike vlogs using a style guide measured from my own Premiere projects: cut lengths, music, captions. They render, export a timeline I can open in Premiere, and check every cut against rules before I see it. First vlog published in October 2026.
Internship watcherPersonal, running twice a day
A small server job that checks the career pages of four companies every morning and evening and sends a push notification when a new internship appears. Built because the Dynatrace posting was only open for about three and a half weeks last year.
MärchenfuchsOwn app, paused
Bedtime stories for children, generated by AI. Flutter and Supabase. A five-stage safety chain checks every story. A test campaign of 50 stories against production found that the database was behind the code, which broke interactive stories.
Training app for phone and watchPersonal, in use
One HTML source becomes a phone app and a native Galaxy Watch app, built without Gradle in about ten seconds. I use it for my five workouts a week.