Philipp Grabner

I build and run software by directing AI coding agents. I write the rules and the spec, split the work, demand proof that it works, and operate what ships. When something breaks, I measure before I fix.

I don’t write most of the code by hand, and I won’t pretend to. What I own is whether it works.

Now
16, student at HTL Klagenfurt Mössingerstraße, electronics and computer engineering, year 3 of a 5-year technical college (graduating 2029)
Looking for
A summer internship in 2027, from 12 July, 4 to 8 weeks, in Klagenfurt or within a daily commute
Based in
Techelsberg am Wörthersee, Carinthia, Austria
Languages
German (native), English (fluent), Spanish (basic)

Selected work

agentfenster

My own product. Windows automation for AI agents, on a desktop you can’t see. Private beta, started August 2026, launch planned for early 2027. agentfenster.com

AI agents that operate desktop apps usually take over your screen and mouse, or run in a cloud VM that doesn’t have your logins. Worse, they report success when a click hit nothing. agentfenster gives the agent its own hidden Windows desktop inside your session: same apps, same logins, and you can watch it live or take over. Claude Code and Codex drive it through an MCP server.

The rule the whole product is built around: never a silent success. After every action it reads the window again. If nothing changed, it says so and stops instead of clicking on.

  1. AgentClaude Code or Codex asks for an action, for example “click Save”.
  2. MCP serverFinds the element in the window’s accessibility tree, not in a screenshot.
  3. Hidden desktopSends the action to the window itself. Your mouse and keyboard are never used.
  4. CheckReads the window again and compares. Did the element react, did the window change?
  5. ReportChanged: names the element it hit. Unchanged: fails loudly and stops.
One action, end to end. The check step is what makes it observable instead of hopeful.
Measured on my laptop while hardening the beta (August to October 2026)
WhatBeforeAfter
Time per click2.5–7.7 s0.1–0.5 s
Opening Edge on the hidden desktop4.9 s0.8 s
CPU with three people watching the live view125 %29 %
MCP server CPU while idle132 %0.6 %
Tokens the agent pays per session before doing anything3,7001,900

CPU is per-core percent, so 125 % means one and a quarter cores.

What I did

  • Came up with the product and its rules, for example “agents must never take over my mouse”, written down as hard project rules every agent reads first.
  • Decided the architecture questions with the agents, including what Windows can’t do (store apps open on the user’s desktop, the clipboard is shared) and stating those limits openly.
  • Used it for my own GUI work every day and logged every failure: about 95 detailed bug reports and 163 run reports so far.
  • Set the bar: a real task has to work end to end, not just a green test suite.

What the agents did

  • Wrote the code: Python, Win32 and UI Automation, a WebView2 interface, the MCP server.
  • Wrote the tests: over 9,500 automated tests (collected on 8 October 2026).
  • Measured with profilers and timers before changing anything, then proved the fix with the same measurement.
  • 2,389 commits between 6 August and 7 October 2026.

Where it went wrong

On 5 October agentfenster reported that a text field “holds the text” while the field was actually empty: exactly the silent success the product exists to prevent. It was caught during daily use, went into the bug log, and the fix landed in the same session with a test that reproduces it. A green suite had not caught it, which is why real tasks are part of the definition of done.

The code is in a private repository because agentfenster is meant to become a paid product. I’m happy to walk through it, the architecture notes and the bug log in a call.

Running a production web app

Operations for a small coaching company (our family business). Live since 15 September 2026, several hundred customers.

The company sells a members’ web app: access codes, checkout, generated PDFs, a journal. I’m responsible for keeping it running: deploys, customer problems, and a support service I set up next to it. The support service sorts incoming requests with one model, drafts replies with a stronger one, and reproduces reported bugs in a separate copy of the app before anything touches production. It has 386 automated tests and an emergency stop: one file on the server switches it off.

Where the AI was wrong

A customer’s app kept crashing on her iPhone. The bug-fixing agent reported “fixed” four times. The usage data said otherwise: she kept bouncing back into the app within 40 seconds of opening one section. So we stopped trusting the agent’s claim and added diagnostic events around that exact screen. As the first fix, the heavy world map on that page now loads on iPhone only when tapped. The lesson: “fixed” needs evidence from real usage, not the agent’s word.

Stack: Node in Docker, SQLite, Caddy, a Python service under systemd, push alerts, automatic failover between AI accounts. I don’t show screenshots or names here because it’s a customer-facing system.

Mitschrift, an AI note-taker for my class

Side project at school. In daily use by two classmates and me since late September 2026.

Every lesson becomes one living note built from photos of the board, typed text, handwriting and, where allowed, audio. Audio is transcribed live in one-minute chunks, a short summary updates every minute, and a final note is written at the end of the lesson. At night my laptop re-transcribes the audio with a local model on the GPU for better accuracy.

Two decisions I made before writing a line of spec: audio is only recorded with the teacher’s permission, which is logged, and raw audio is deleted after seven days.

A problem worth knowing about: speech recognition models hallucinate on classroom noise. Ours kept producing the sentence “Untertitelung des ZDF” (subtitles of a German TV channel) when nobody said it, because it learned that phrase from TV subtitles. Transcripts are now checked against a list of these known phrases, and English filler from YouTube outros, before anything reaches the note.

Stack: FastAPI and SQLite, a Svelte web app that installs like an app, an Android app with its own recording service. 57 commits since 22 September 2026. Planned entry for Jugend Innovativ 2026/27.

Measured, then fixed

Four real problems from my own machines, written up like incident traces. In each one the agents measured first and changed things second.

  • Symptom
  • Measurement
  • Fix
  • Fix that didn’t work
  • Result

The laptop froze for seconds at a time

9 to 15 August 2026

  1. Symptom

    Mouse and keyboard freeze for seconds, video editing looks like it crashes. Task Manager shows nothing unusual: no process is using the time.

  2. Measure

    The Windows hardware error log (WHEA) has 324 corrected PCIe errors in 14 days, all from the same PCIe port. The time is going into interrupt handling, which Task Manager doesn’t attribute to any process.

  3. Fix

    Replace the 2017 driver of the Realtek network chip on that port and switch off all its power saving. Next day: 845 errors in 30 seconds. Didn’t work.

  4. Fix

    Switch off PCIe link power management (ASPM). Errors continue, all “receiver errors”: the link is electrically marginal and the OS can’t heal it.

  5. Fix

    Hand PCIe error handling to the firmware with bcdedit /set {current} pciexpress forcedisable, so the OS no longer drowns in error interrupts.

  6. Result

    0 hardware errors over the next four days. A “crash” reported on day six turned out to have no crash artifact at all: it was Windows Update and the search indexer catching up.

What I took from it: before treating something as a crash, check whether a crash artifact exists.

Crashes that looked like a memory leak

4 to 7 August 2026

  1. Symptom

    The laptop crashes eight minutes after boot. The suspicion is a memory leak.

  2. Measure

    No leak. A tool restored 13 agent sessions at once, 16 in total. Each started about 70 helper processes (MCP servers) and committed about 5 GB: 81 of 101 GB committed memory.

  3. Fix

    Cut the default helper servers from 18 to 8, the full set loads on demand. 5.1 to 3.24 GB per session, measured. Added a watchdog that cleans up orphaned processes.

  4. Symptom

    The next evening the watchdog fires at 74 % memory. It can’t fix it.

  5. Measure

    Ten restored sessions had not written a single byte in 12 hours but still held about 800 MB each. Idle sessions aren’t orphans, so the watchdog ignored them. A note-taking helper also launched up to 15 one-question AI calls per run, each starting the full set of 19 helper processes.

  6. Fix

    The watchdog now ends sessions idle for 90 minutes together with their whole process tree, and one-question calls start with no helpers at all: 1 extra process instead of 19.

  7. Result

    Committed memory 75 to 53 GB that evening. Two days later the same pattern: 195 idle processes removed, 95 % to 80 %.

What I took from it: “out of memory” was a fan-out problem, not a leak. The fix that mattered was in the rule for what counts as idle, not in the amount of RAM.

Python at 3,000 % CPU

17 September 2026, about one hour

  1. Symptom

    During a text-recognition batch job the laptop is close to unresponsive. Python processes use about 3,000 % CPU.

  2. Measure

    Ten worker processes at about 300 % each. The OCR library runs on onnxruntime, which by default starts a thread per core in every worker and ignores the usual OMP_NUM_THREADS setting. Ten workers times every core is oversubscription.

  3. Fix

    One thread per worker, at most six workers, lower process priority. Measured detail: child processes did not inherit the lower priority (parent “below normal”, workers “normal”), so it’s now set inside each worker.

  4. Result

    The job runs in the background while the laptop stays usable. The rule now applies to every parallel ML job: set threads per worker explicitly.

Console windows popping up on my screen

6 September 2026

  1. Symptom

    For about ten minutes, terminal windows keep popping up on my screen while agents work in the background.

  2. Measure

    A window watcher plus a controlled test: start a plain console directly on agentfenster’s hidden desktop and count what appears on mine. 2 windows. Windows 11 hands every new console to Windows Terminal, whose single process lives on my desktop.

  3. Fix

    Default terminal back to the classic console host, plus one launcher that starts scheduled tasks without a window. Every change reversible, with backups.

  4. Result

    Same test, 0 windows.

Bars show the order of the work. Where there is a date axis, positions are approximate within a day.

How I work

My projects are written by AI coding agents, mostly Claude Code, often several in parallel. That only works with structure. This is the loop I run.

  1. Write the rules before the code. Each project has a rules file that every agent reads first: what the product promises, what must never happen, and the limits we have already measured. agentfenster’s starts with seven rules, the first one being “never a silent success”.
  2. Split the work. Parallel agents get separate work packages and separate copies of the repository, so they don’t overwrite each other. I decide the order and what waits.
  3. Demand evidence. A change isn’t done because the agent says so. It needs tests, a measurement, or a real run, and a skipped test counts as a finding, not as green.
  4. Use it myself. I run my own tools daily. Every failure goes into a log with the same fields: expected, happened, workaround, fix, status.
  5. Operate it. Linux servers, Docker, systemd, cron jobs, Cloudflare, push alerts to my phone when something needs me. When it breaks, back to step one with what we learned.

As long as that doesn’t work yet, it isn’t a mature product.

Me, in August 2026, about a real end-to-end task agentfenster had to pass, while thousands of tests were already green

What I can do

  • Plan software and break it into work an agent can verify
  • Review changes, read logs, measure and debug on Windows and Linux
  • Run servers and deployments (Linux, Docker, Caddy, Cloudflare)
  • Electronics from school: circuits, lab measurements, Multisim
  • Video: filming, drone, editing (Premiere Pro)

What I can’t do yet

  • Write Python or TypeScript by hand. I’m learning to read the code I ship.
  • Java at school level, my third year of it.
  • Large-scale systems with real traffic. That’s exactly what I’d like to learn in an internship.

Also built

  • Video editing pipelinePersonal, in use

    Agents cut my mountain-bike vlogs using a style guide measured from my own Premiere projects: cut lengths, music, captions. They render, export a timeline I can open in Premiere, and check every cut against rules before I see it. First vlog published in October 2026.

  • Internship watcherPersonal, running twice a day

    A small server job that checks the career pages of four companies every morning and evening and sends a push notification when a new internship appears. Built because the Dynatrace posting was only open for about three and a half weeks last year.

  • MärchenfuchsOwn app, paused

    Bedtime stories for children, generated by AI. Flutter and Supabase. A five-stage safety chain checks every story. A test campaign of 50 stories against production found that the database was behind the code, which broke interactive stories.

  • Training app for phone and watchPersonal, in use

    One HTML source becomes a phone app and a native Galaxy Watch app, built without Gradle in about ten seconds. I use it for my five workouts a week.

  • Orbit LeapPublic code

    A one-tap browser game built in a day. Play it.

  • Claude’s GamePublic code

    A game built by an AI one YouTube comment at a time: the top comment decides the next feature.

Background

School and awards

  • Since 2024: HTL Klagenfurt Mössingerstraße, electronics and computer engineering
  • Year 1 (2024/25) passed with distinction
  • Technicus Award 2024, 1st place: a scale model of an energy store that uses potential energy
  • Technicus Award 2023, 3rd place, with the project “Kann ich mir selbst mit Minecraft Java beibringen?”
  • Känguru der Mathematik, 2nd place in Carinthia

Outside of school

  • Mountain-bike YouTube channel since 2022, over 1,000 subscribers: filming, drone, editing
  • Bikeventure, a yearly short film with friends. 2026: nine cameras, four shooting days
  • FPV drones, 3D printing, CAD in Fusion 360
  • Training five times a week, skiing in winter

Log

  • This page goes online.

  • agentfenster: clicks 5 to 21 times faster, idle CPU of the MCP server from 132 % to 0.6 %.

  • First vlog cut with my video pipeline published.

  • iPhone crash in the coaching app found through usage data after four false “fixed” reports.

  • Mitschrift started, in class use a week later.

  • Coaching web app launched, several hundred customers.

  • First commit of agentfenster, during the BridgeMind Vibeathon.

Contact

philipp@philipp-grabner.com

I answer within a day. If you’d like to see agentfenster or the code behind any of this, I’m glad to show it in a call or in person in Klagenfurt.