
(Just looking at my game's current demo screen feels embarrassingly dated.)
This month was busy with business meetings. I met with four clients in total, and one of them led to a deal. The meeting was about delivering an in-house ERP solution for a startup, specifically a request to add RAG functionality.
Between preparing for meetings, I had no time to write or pursue any hobbies.
With Korean holidays mixed in and quite a few business meetings, I interviewed with about three clients and completed a delivery project for one of them. The total came to roughly 450. Compared to the old days when there were plenty of 800-range jobs, I can't help but feel how much the rates have dropped. But these days, everywhere you go someone claims they can build it with AI, so quoting high just sends clients to another vendor, making it hard to price work at a premium. It feels like there's a psychological anchor line, and that line tends to sit around 450 per month.
I also heard that a major Korean conglomerate recently had a serious problem related to vibe coding. Apparently management gathered the PMs and boasted they would eliminate all external development contracts by coding everything through prompt input alone, but word on the ground is that they still can't resolve the bugs. A senior colleague told me things should normalize in about two to three years. (I know which company it is, but I won't say.)
My senior colleague says it's because they didn't develop a proper harness, but I don't see it that way.
Anyway, since I had some free time this month and the topic of harnesses came up in conversations in Korea, I want to write down a few thoughts.
I think the idea that every developer should have their own harness is meaningless.
Engineer your own harness, burn through cheap API calls, and drive down unit costs? Honestly, I'm not sure this is worth anything beyond a toy relative to my time. I'm not someone who builds AI frameworks for a living; my actual work is designing and delivering ERP systems and websites.
So the idea that I should hand-code a harness tailored to my own workflow feels pointless to me. Building RAG and similar things from scratch and trying them out leads to the same conclusion. The supposed advantage of DIY tools like pi.dev is that you can fully personalize them to your workflow, but the critical drawback is that you get no community support. Official tools like Codex or Claude Code, on the other hand, give you access to a large community and plenty of references whenever something goes wrong.
In the first place, AI model APIs themselves are increasingly evolving to be optimized for their own official harnesses. Frontier models will inevitably end up tightly coupled with their native harnesses. Using a custom harness means you have to manually craft every tool description and prompt, and each time the model updates you have to remove tools that worked well before or rework the entire configuration.
People say "the API is cheap, so that's the advantage!" but in reality my time (labor cost) is far more expensive. I have a primary job that needs to get done. Running heavy multi-agent setups with LlamaIndex or LangGraph isn't the answer either, since you have to reconfigure everything whenever the model changes. And that's work for a company with an in-house development team, not for a solo freelancer like me. That's overengineering.
To be fair, pi.dev is still useful for certain tasks. When Claude Code or Codex raises security concerns and I don't want to run a proxy through something like GLM, I'll just fire up pi.dev. But in that case, I find there's little meaningful difference between that and attaching tools via MCP or Skills on top of an official harness, assuming you're running a subscription anyway.
I'll grant that DeepSeek writes code well, but when it comes to reasoning about a large codebase it still struggles. The moment you start building tools and pipelines just to help the model understand a big codebase, you're burning through your own time again.
What frustrates me most is that after spending all that time building tools, a model upgrade renders some of those tools unnecessary or forces you to rip them out entirely. The prompts and tool descriptions I optimized in the past actually seem to interfere with the improved capabilities of the new model. In other words, the features of the harness I built depreciate over time.
On top of that, because AI is inherently non-deterministic and probabilistic, whenever an error occurs it's often hard to tell whether the problem is in my custom harness or just the model hallucinating. It's debugging hell.
So my conclusion is simple. The best approach is to run frontier models on their official subscription plans, and to tunnel through that major harness to connect a custom provider (such as OpenRouter) as a proxy.
These days I set up custom providers for Codex or Claude Code and pull in other AI models through them. I find myself reaching for tools like pi.dev less and less. At this point, their only remaining advantages are hobbyist use or the ability to run in a fully air-gapped environment; beyond that, they feel like DIY toys that demand you pay the management cost in your own precious time.
A note on AI's limitations
Lately, every field I visit seems to be embracing AI workflows and AI-as-a-cure-all thinking. Honestly, since software engineering hasn't properly taken root in Korea, I tend not to pay close attention to the discourse coming from that side. I mostly talk with Chinese and American developers. (I recently heard from some American developers that a rumor is circulating about a Notion-like feature being added to Claude.)
Personally, despite all the AI-as-a-cure-all talk, the more I keep using frontier AI models, the more something feels off. The models are clearly progressing, yet as a developer the change I actually feel is no longer as significant as it used to be.
This is especially true in coding. Even at around GPT 5.6 Sol, the models were already strong enough for function- or class-level implementation. Models since then are faster, accomplish the same tasks with fewer tokens, and have noticeably improved in areas like image and 3D generation, but the quality of the code itself has actually declined somewhat.
In other words, it sometimes feels as though gaining one capability comes at the cost of gradually sacrificing another.
Astra is impressive for things like 3D asset generation, but for the coding I actually do, I haven't found it to be better than 5.6 Sol. The recent OpenAI announcements also read, from a developer's perspective, more as stories about usage adjustments, pricier premium tiers, and new product interfaces than as any fundamental leap in model capability.
Current models can already generate small units of code with considerable precision. Simply nudging that accuracy a little higher is unlikely to produce a significant change. Moving to the next level requires one of two directions, it seems.
One is for the model to genuinely understand and work within a massive codebase of millions or tens of millions of lines.
The other is for intelligence itself to advance to the point where it can compress and solve enormous problems with very little output.
The latter is unrealistic given statistical and computational constraints, while the former is possible but prohibitively expensive. My sense is that AI currently solves neither, and is instead papering over the gap with harnesses as a stopgap measure.
This perspective aligns reasonably well with the problems I actually encounter in day-to-day coding.
The area where AI is weak right now is global consistency.
Function implementation
Method extraction
Writing generics
Result/Option chaining
Policy table
lookup table
RefactoringAt the granular level, AI already outperforms humans. Overall, there hasn't been much meaningful difference since around GPT 5.2.
The real challenge is maintaining the relationships of the overall system that those functions belong to over the long term.
As a codebase grows, what matters is not how well new code is written but rather
AI produces good implementations far faster than humans when looking at the code right in front of it, but when given a long task it often keeps patching local solutions and ends up disrupting the overall architecture.
That's why, when you commonly use a /goal command to give AI a target and let it code away, the codebase frequently ends up a mess.
Even in my own AI-assisted work, this problem is exactly why I don't hand over the entire task; instead I look up GitHub references myself and build by following those structures. Microsoft's GitHub and dev samples tend to provide plenty of such examples.
I first lock down the spec, create an MVP and vertical slice, then break the work into function- or class-level units and delegate the implementation to AI. In a large codebase, I have AI investigate the existing code and dependencies first, verify the results myself, and then re-specify the exact scope of work.
In this workflow, the difference between the latest model and a slightly older frontier model turns out to be smaller than you'd expect.
That's because the ability to write a single function well has already reached quite a high level.
My feeling is that AI will stay in this range for a long time.
Most demo videos will probably keep competing on how visually impressive a one-shot result can look.
A single function, a UI screen, or a 3D asset is self-contained data, so AI can easily mimic the pattern. But the global consistency of a system is determined by business rules that exist outside the text of the code and by decisions made in offline meetings. AI has never learned this "invisible context," so as code grows longer it gets increasingly buried in the text right in front of it. The conclusion, then, is that the competition will come down to how visually stunning a result an AI demo can produce.
I want to call this demo-driven capitalism.
Anyway, work has been so heavy lately that I haven't had time to research or read papers. For me personally that's a good thing, but at the same time I keep thinking that I need to build my own product to escape this grind. It's always hard.