Blog

Notes on building AI agents that hold up

Field notes from building aibuddy in the open — the runtime, context engineering, and the coordination problems a bigger model doesn't solve for you.

Topic
Source
Company
Sep 4, 2026 15 min read Zhang Xiaojun

Models Are Not the End: Agents and the Next Organizational Shift

The agent ecosystem is still waiting for an entry point that dramatically lowers the barrier to use. Models will absorb simple scaffolding, but not complex work—and the next advantage will come from task loops and intelligence flywheels.

Engineering
Aug 30, 2026 9 min read Original

Why AI agents need a harness, not just a better model

Better models raise the ceiling, but reliable agents come from the system around the model: context, tools, constraints, verification, correction, and observable loops.

Company
Aug 29, 2026 10 min read Original

AI-native companies reorganize context

AI-native transformation is not about adding AI to old workflows. It is about reorganizing company information so agents receive the right context before each decision and return evidence for the next one.

Research
Aug 28, 2026 11 min read OpenAI

AI Jobs Transition (1): Capability is not displacement

How much work AI can perform tells us where change may begin, not whether a job will disappear, reorganize, or grow. We need a better framework than another automation-risk ranking.

Research
Aug 28, 2026 8 min read OpenAI

AI Jobs Transition (2): The work that still needs a person

A job can be highly exposed to AI and still require a person for accountability, trust, or physical execution. That protects the role from full automation, but not from redesign.

Research
Aug 28, 2026 8 min read OpenAI

AI Jobs Transition (3): Cheaper work can create more work

AI reduces the labor needed for each unit of output, but lower costs can also unlock new customers and new demand. Employment depends on which force wins.

Research
Aug 28, 2026 8 min read OpenAI

AI Jobs Transition (4): The capability gap is an operating problem

Models can already affect far more work than organizations actually delegate to them. Closing that gap requires context, permissions, evaluation, and workflow redesign—not more prompt tips.

Research
Aug 28, 2026 8 min read OpenAI

AI Jobs Transition (5): Four paths, not one future

The labor market is not moving toward one AI outcome. Four transition paths require different company decisions, worker strategies, and policy responses.

Research
Aug 28, 2026 9 min read OpenAI

AI Jobs Transition (6): Turn the framework into an operating plan

The useful output of AI labor research is not a prediction. It is a repeatable way to choose what to automate, what to redesign, where to grow, and what to monitor.

Research
Aug 28, 2026 15 min read Original

Why enterprises need AI

Much of enterprise work depends on language, unstructured information, and professional judgment that traditional software cannot economically encode. Foundation models are turning that work into a general, scalable software capability.

Engineering
Aug 24, 2026 7 min read OpenAI

Harness Engineering (1): When engineers start designing the agent's environment

When code generation is no longer scarce, engineering shifts from writing every implementation to designing environments, expressing intent, and building feedback loops.

Engineering
Aug 24, 2026 7 min read OpenAI

Harness Engineering (2): Give the agent a map of the repository

A giant AGENTS.md does not solve context. Agents need a short entry point, layered knowledge, and a system of record that stays aligned with the code.

Engineering
Aug 24, 2026 6 min read OpenAI

Harness Engineering (3): Make the running system legible to the agent

An agent that can edit code still cannot verify the product. Autonomous engineering requires applications, browsers, logs, metrics, and acceptance criteria that agents can inspect directly.

Engineering
Aug 24, 2026 5 min read OpenAI

Harness Engineering (4): Turn engineering rules into executable constraints

Agents do not follow a rule forever because they read it once. A reliable harness turns architecture, quality requirements, and engineering taste into checks the system can enforce.

Engineering
Aug 24, 2026 6 min read OpenAI

Harness Engineering (5): When agent output exceeds human attention

Once agents increase code throughput, one-by-one human review becomes the next bottleneck. Checks, review, merge policy, and recovery must be configured by risk.

Engineering
Aug 24, 2026 7 min read OpenAI

Harness Engineering (6): Agents copy technical debt too

Agents learn from the patterns already in the repository, including duplication, drift, and temporary patches. High-throughput systems need continuous quality garbage collection.

Community
Aug 12, 2026 7 min read Original

AI Builders: From Using AI to Shipping Systems

AI builders do more than use tools well. They put AI into real settings and turn it into systems that can be delivered, verified, and improved.

Research
Aug 11, 2026 12 min read Original

Ideas are cheap. Execution is expensive.

In the AI era, many directions become obvious. The real advantage is not having the idea first, but defining the problem, filtering noise, building a verifiable execution system, and carrying long-horizon work to completion.

Company
Aug 10, 2026 7 min read Anthropic

AI-Native Startup (1): What Actually Gets Rebooted

AI does not make startups easy. It compresses the loop from idea to evidence to product to feedback.

Company
Aug 10, 2026 9 min read Anthropic

AI-Native Startup (2): Existing Company Data Is the First Asset

Before an AI-native company asks AI to write more code, it has to turn customer, product, sales, and ops data into reusable context.

Company
Aug 10, 2026 7 min read Anthropic

AI-Native Startup (3): The Founder Becomes an Orchestrator

AI does not make the founder's job lighter. It moves the work from direct execution to system design, judgment, and orchestration.

Company
Aug 10, 2026 8 min read Anthropic

AI-Native Startup (4): In the Idea Stage, Building Is Not Validation

AI makes building dangerously easy, so the Idea stage has to prove the problem is real, specific, frequent, and worth solving.

Company
Aug 10, 2026 8 min read Anthropic

AI-Native Startup (5): In the MVP Stage, Boundaries Matter More Than Code

The first artifact of an AI-native MVP should not be code. It should be scope, architecture, metrics, and context.

Company
Aug 10, 2026 8 min read Anthropic

AI-Native Startup (6): Launch Is Where Product Becomes Company

Launch is not the announcement. It is the stage where founder improvisation becomes a repeatable operating system.

Company
Aug 10, 2026 8 min read Anthropic

AI-Native Startup (7): The Real Moat for AI Products

The moat is not the model. It is domain knowledge, user behavior data, workflow embeddedness, and the operating loop that compounds them.

Company
Aug 10, 2026 7 min read Anthropic

AI-Native Startup (8): Same Founder Job, New Rules

AI does not replace founder judgment. It makes judgment, context, and learning speed more central than ever.

Engineering
Aug 8, 2026 15 min read Original

From using AI to becoming an AI-native team

AI-native teams are not defined by how many people use chat. They are defined by whether agents can participate in real work with shared context, scoped authority, verification, and accountable humans.

Company
Aug 7, 2026 12 min read YC

Build your company as an intelligence layer, not an org chart

AI shouldn't be a tool your company uses — it should be the operating system your company runs on, and that rewrites the org chart, not just the output.

Company
Aug 6, 2026 13 min read YC

How to build an AI-native services company

Some of the biggest companies of the next decade won't sell software — they'll be law firms, insurers, and tax practices rebuilt from scratch with AI doing most of the work.

Company
Aug 5, 2026 11 min read YC

How to pick a startup idea

The perfect idea doesn't exist in the abstract — the only way to find what works is to pick one, burn the other boats, and go deep enough to run your customer's business.

Company
Aug 4, 2026 12 min read YC

How to get your first ten customers

Your first ten customers almost never come from a tool — they come from your network, showing up in person, and a willingness to do the things that don't scale.

Research
Jan 14, 2026 10 min read Original

Context engineering is the bottleneck in long agent runs

Long agent runs fail when the model sees the wrong slice of state. The hard part is not stuffing the window; it is managing context as working memory across turns, tools, compaction, memory, and cache.