Hacker News
an hour ago by Garlef

I think they even more so need deterministic feedback:

I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.

https://habit-hooks.com/

I'm using it to foster IOSP (integration operation segregation principle) for example.

5 hours ago by spike021

I think whichever one is used, there needs to be a way to enforce what's written.

If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.

5 hours ago by zahrevsky

There is a command in oh-my-pi called "/omfg <problem>". You explain what is wrong with agent's response, and it writes a hook to make sure that the problem doesn't happen again. It then re-runs your previous prompt to make sure that hook is triggered, and if not, it rewrites the hook to make your previous prompt trigger the hook. Then each next agent's response is checked by the hook, and if it is triggered, the agent receives feedback on what's wrong and what must be done differently.

2 hours ago by aliasxneo

I've written a lot of custom hooks this way. It's an amazing feature.

4 hours ago by jkhdigital

Treat it like any other software system: rules that must not be violated are enforced by static type-checking or a trusted runtime monitor. There’s no other option.

3 hours ago by koolba

Except you can’t do that unless the runtime itself can reason about what’s being executed.

Otherwise you can get a python one liner that execs a different script engine.

5 hours ago by dboreham

I think it's kind of cute the way it writes Python scripts, but I've never seen it do that when the relevant native tool is on the PATH. It's like the most competent ever intern, on speed. No tool to convert SVG to PNG? No problem, I'll write a Python program to do that!

2 hours ago by sick_of_slop

[dead]

7 hours ago by bushido

Something I started doing recently was writing out principles instead of memories.

Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.

I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.

I've had surprisingly good adherence from agents on this technique.

[0] https://principledriven.dev/

7 hours ago by skybrian

It sounds interesting, but this looks too much like a marketing website (huge text everywhere) and not enough like a documentation website, so it was hard for me to see how it might work.

6 hours ago by bushido

Fair feedback. I admittedly didn't spend enough time making it simpler.

The github repo might be more helpful: https://github.com/Principle-Driven/pdd

4 hours ago by skybrian

That's better.

I was wondering about this bit:

> "Code comments cite a versioned token where code depends on the rule."

What does a token look like? I went looking for an example, but I don't know what I'm looking for.

8 hours ago by gregwebs

Agreed, and this seems better.

My thought though has always been that I don't want there to be agent-only designated documentation.

I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.

6 hours ago by cyanydeez

I've got basically a loop of docs, test (TDD) and code. Starting with and IMPLEMENT-<plan name>.md. I ask to revise as TDD, then loop.

Once qwen3.8-flash-next showed up, it cañ go "forever" with dynamic context pruning.

It's fascinating for local coding.

6 hours ago by j45

A lot of agentic software development feels like cowboy coding, set it up, let it rip and see what it figures out.

The slightest amount of guidance, input from experience can make a huge difference.

7 hours ago by isaachinman

I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects

https://github.com/isaachinman/encephalon

4 hours ago by jmtulloss

Evals or it didn’t happen.

Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.

2 hours ago by DriverDaily

The brain has mechanisms that can organize experiences without requiring that relationship to be expressed as a sentence.

Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.

Documents can’t be queried efficiently like that, you need a database.

an hour ago by apsurd

I get what you're saying but it seems presumptions to compare a graph database to how our brains work. And also that LLMs <-> (the way they do memory) is the right analog to humans <-> memory.

I take the article's point more directly. It's just a straighter line to have clear communication through documentation than to fuddle around with the perfect memory setup.

8 hours ago by alienbaby

I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..

however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.

Daily Digest

Get a daily email with the the top stories from Hacker News. No spam, unsubscribe at any time.