Last updated: 2026-10-05
From Semantic Search Towards Cognitive Agency: The Accidental Architecture of Memsearch
A personal knowledge tool that acquired several of the functions a cognitive architecture needs, and what that does and does not show
Memsearch did not begin as an attempt to build an agent. It began as a way of finding material across one person's working files and conversation histories, and it is still, at its core, a retrieval tool. As functions were added, though, its design started to resemble the cognitive cycle this module uses as its shared diagram. It senses its environment, keeps several kinds of memory, abstracts patterns across them, proposes connections, checks some of those connections against outside sources, revises them, and changes its own workspace.the diagram is a classic control loop
This page treats memsearch as a case study rather than a showcase. It is a work in progress. Its idea-synthesis pipeline is experimental, and the project's own documentation says that its output is speculative machine synthesis, to be investigated rather than cited. Nothing here claims that the system has been validated, and it is not yet a fully autonomous agent. The page asks where a search tool becomes a cognitive system, and where a cognitive system becomes an agent, and it uses the project's real behaviour to make both questions concrete.
What Memsearch Currently Does FoundationalKnowledge that endures for decades — core principles
Memsearch chunks and embeds the files of several tracked project directories, every conversation transcript from the author's coding assistant, including those from sub-agent runs, and a local library of PDFs. It stores the embeddings in a local vector database and builds a knowledge graph over them, including a map of nearest neighbours and two kinds of gap. Interpolation gaps are pairs of nodes that are related in embedding space but have nothing connecting them. Extrapolation gaps come from each cluster's outer edge: a small region spanned by several of its most outlying members reaches out in several directions, and the nearest existing node in each direction is a candidate for a bridge. A local web interface shows the graph and a board of candidate ideas and projects.
On top of the retrieval layer sits an idea-synthesis pipeline with seven commands. synthesize proposes cross-domain ideas from the graph's interpolation gaps, and synthesize-region explores the extrapolation regions in one of three modes: point, which treats each sampled direction's nearest neighbour as an ordinary connection; region, which describes the whole region without a concrete anchor; and combined, which grounds the region description in those nearest neighbours. Only point mode has been run end to end so far. verify takes a report's open questions to a headless instance of a hosted model with web search, and writes the findings back into the report. refine revises an idea against those findings. synergy looks for shared techniques or dependencies between live ideas. consolidate rolls related ideas into a named parent project, and decompose breaks a large idea into sub-projects with stated interfaces. Reports produced by synthesis are mined back into the store, so they become part of the material that later cycles search and draw on.
Three Tiers of Material Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The material the system searches is not uniform, and the tiers the site uses for knowledge decay[the half-life of knowledge] give a useful way to see the difference. Conversation transcripts are ephemeral. Each records a particular exchange, usually about a particular problem, and it is rarely consulted again once the problem is solved. The project artefacts the system indexes, such as specifications, papers, code, and notes, are closer to the foundational tier, because they keep their value over years. The method itself, meaning the way the pipeline is organised and the reasons for its design, sits at the applied level. It is used in practice and revised as it is used, but it is not yet settled.
The tiers matter for the system's behaviour. A search that returns a transcript from two years ago should be read differently from one that returns a specification, and the store does not currently mark the difference. The ephemeral-to-foundational pathway material describes how a quickly acquired idea can be promoted to a durable one, and memsearch does not yet do that promotion. It indexes both kinds of material with equal weight.weight is not relevance
A Nightly Cycle Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A scheduled task runs memsearch once a day at 03:00. Its job re-mines the tracked project directories and the conversation transcripts, fetches new external sources for the PDF library, removes the chunks of any file that no longer exists on disk, and rebuilds the knowledge graph. Once the indexing and graph rebuild are finished, the same run goes on to synthesis. It proposes ideas from the graph's gaps, refines each one with the local model, mines the finished report back into the store, and scans the ideas for synergies. The overnight work is therefore both maintenance of the representation and generation, and the generated material is written back into the store that the next night's run will index.
The time was chosen for convenience, since a job that runs while the machine is idle does not disturb daytime use. The schedule nonetheless produces a resemblance to the sleep-based consolidation described on the sleep consolidation page. The system takes in experience during the day, in the form of files written and conversations held. It reorganises that material overnight, extracting the patterns that the embeddings and graph carry, and it removes material that is no longer there. The parallel with the biology is partial, and it should be read as a resemblance of timing and function rather than of mechanism. Nothing is replayed, and nothing is retained because it was reinforced. The pruning removes sources that have been deleted, not weakly connected material, so it is closer to clearing away the leftovers of a day than to the synaptic downscaling that Tononi and Cirelli propose. The parallel is this page's own reading, not something the sleep page claims about software.cf. biological consolidation via replay, not just timing
Mapping the System onto the Cognitive Cycle Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The table below is an interpretation. It reads memsearch's commands against the stages the cognitive-cycle page sets out[the cognitive cycle]. The project was not designed around that model, so the correspondences are the author's reading of the code, and the last column says how strong each one is.
| Cognitive function | Memsearch mechanism | Interpretation |
|---|---|---|
| Sense | mine, mine-convos, and source ingestion | Acquiring material from an environment made of project files, conversations, and documents |
| Perception | Chunking, parsing, and metadata extraction | Turning source material into units the system can process |
| Short-term context | The current query, the records it selects, and the candidate idea under examination | The immediate working situation of one operation |
| Episodic memory | Stored project files and conversation histories | Records of particular past work and interactions, kept with their source |
| Abstraction | Embeddings, clusters, and nearest-neighbour relations | Similarity structure drawn across many records |
| Representation | The vector store and the derived knowledge graph | The system's current organised account of the user's material |
| Think | Semantic retrieval and synthesis over selected material | Finding relations relevant to a query |
| Imagine | Candidate ideas proposed from gaps between clusters | Possibilities not explicitly present in the source material |
| Predict | No dedicated stage | Consequences of a candidate idea are not modelled explicitly, so this stage is largely absent |
| Verify | verify, which searches the web for the open questions | Testing an internally generated idea against sources outside the system |
| Reflect and revise | refine | Changing a proposal in response to what verification found |
| Coordinate | synergy and the reconciliation of several decompositions in decompose | Relating independently generated analyses to one another |
| Act | consolidate, decompose, and the idea board | Changing the persistent workspace, not only returning text |
The Predict row is the one most likely to be over-read. A reader might take the decomposition step, which states interfaces and dependencies, for a prediction of consequences. It is closer to a plan for a project than to a forecast of what the project will do. For the purposes of this module, the gap is that memsearch has no stage that runs a candidate forward and compares outcomes.plans are not predictions missing the simulation step
Not a Pipeline Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The commands are usually run in a sequence, which makes it tempting to describe the system as a pipeline: mine, search, synthesize, verify, refine, act. That sequence is a useful way to operate the tool. It is a poor account of what the system is doing, because the important feature is the feedback. Once a generated report or a project definition becomes part of the user's files, a later mining run ingests it, so action changes the environment that will be sensed next.
Environment (project files, conversations, PDFs)
|
v
Sense and ingest
|
v
Persistent memory and representation (vector store, graph)
|
v
Retrieve, associate, and synthesise
|
v
Generate candidate idea
|
+------> external verification
| |
| v
+----------- revision and refinement
|
v
project or knowledge action
|
v
changed environment --> mine again
This is closer to the cycle than a conventional retrieval application is. The cognitive-cycle page makes the same point about biological and artificial agents: the stages are densely connected, action feeds back into the environment, and plans are revised by results. The loop in memsearch also carries a risk, discussed below, that the same feedback can amplify a weak idea.the classic control loop
Not One Memory but Several Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A vector store is often described loosely as an agent's memory. Memsearch shows why that description hides more than it reveals. It holds several things that perform different memory functions.
- Source memory. The original files and transcripts remain the primary record. They preserve particular events and discussions, which is closer to episodic memory than to abstraction.
- Semantic or abstract memory. Embeddings and clusters discard much of a source's exact form and keep patterns of similarity. They work more like abstraction than recollection.
- Relational memory. The knowledge graph records connections between items that may come from different projects and conversations.
- Working context. A query, a candidate idea, a verification report, or a decomposition attempt forms a temporary working set for one operation.
- Prospective memory. The idea board records matters intended for later attention. It stores intentions and unfinished possibilities, which are not memories of the past.
The cognitive-cycle page separates episodic experiences from generalised abstractions and from the immediate working situation. Memsearch makes the same separation in its storage, and the distinction matters because each kind is revised in a different way. Source material should not change because a report about it was written. Generalisations change when the material changes. Prospective entries change when a person decides they are finished.raw logs vs. learned rules vs. current plan
Gap-Finding as Constructive Imagination Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The most interesting feature is the attempt to generate ideas from gaps between clusters. The project's own description of an interpolation gap is precise: a pair of nodes that are related in embedding space but have nothing bridging them. Extrapolation gaps are a different kind of candidate. They lie past a cluster's edge, along directions the cluster's outermost members point toward, so they are regions to explore rather than pairs to connect. That is a geometric fact about the representation, and it says nothing about whether the two areas should be connected.
An embedding-space gap is therefore not an undiscovered fact, a valid research gap, evidence that two concepts belong together, or proof of novelty. The system does not discover ideas latent in the vector space in the way a search engine finds matching records. It constructs candidate relations from the geometry of its representation. Those candidates may be productive, trivial, mistaken, or artefacts of the embedding model. Until they are investigated separately, their status is imaginative.
This fits the cognitive-cycle account of imagination, which treats it as recombining fragments of what is already known into situations that have not occurred, rather than extrapolating the present forward. It also fits the project's own warning that generated connections are leads for human attention, not conclusions.recombine, don't just predict
Verification Outside the Generating Loop Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The robust-agent page argues that verification must come from a path that does not run through the component being verified[designing auditable agentic systems]. Memsearch applies this in a concrete form. The local model that writes synthesis reports has no web access, and it is instructed to name anything it cannot check as an open question rather than assert it either way. A separate step takes those questions to a hosted model with web search, and writes the findings back into the report. The generating stage and the checking stage are separate processes with different capabilities, and that separation is the design point.
Separation of this kind does not make verification independent. Two models can share training sources, assumptions, and failure patterns, and a checking stage can accept the framing that the first stage supplied. Verification in this system should therefore be split into distinct checks, each with its own question.
| Check | Question |
|---|---|
| Source verification | Do supporting sources exist? |
| Claim verification | Do those sources actually support the claim? |
| Novelty verification | Is the alleged connection already well established? |
| Contradiction search | What evidence counts against the proposal? |
| Provenance check | Which parts came from source material, and which were generated? |
| Human judgement | Is the idea worth pursuing in this context? |
The system performs the first and some of the second, and the last belongs to the person. The middle checks are where the design is weakest, because a web search can find a source that loosely resembles a claim without supporting it.
From Retrieval to Agency Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A program does not become an agent because it calls a language model. The stronger agentic characteristics appear when a program maintains persistent state, works toward an object or goal, chooses among possible actions, observes the effects of those actions, updates its model, and continues the cycle. Memsearch implements several of these, but the high-level sequencing is still started by a person, one command at a time.
| Level | Capability | Memsearch example |
|---|---|---|
| Retrieval system | Finds relevant stored material | Semantic search |
| Representational system | Organises relations among stored material | Embedding space, graph, and clusters |
| Cognitive support system | Generates and evaluates possible connections | Synthesis, verification, and refinement |
| Agentic workflow | Selects and performs operations that change persistent state | Consolidation, decomposition, and project organisation |
| Autonomous agent | Chooses its own goals and continues acting without prompting | Not present; not needed for the tool to be useful |
Memsearch sits in the agentic-workflow row. It is a cognitive tool with agentic workflows rather than an autonomous agent. That intermediate category is both accurate and useful for teaching, because it shows which functions can be built incrementally, and which would need a different kind of control to be trusted.
Human and System Roles Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The tool is not simply automating the production of ideas. Its division of labour looks like this.
| System contribution | Human contribution |
|---|---|
| Searches across more material than anyone can hold in working memory | Recognises which findings are significant |
| Detects weak or distant similarity | Judges whether the similarity is meaningful |
| Generates possible bridges between areas | Supplies disciplinary and personal context |
| Repeats verification and revision | Decides when the evidence is adequate |
| Tracks candidate ideas and their relations | Chooses goals and priorities |
| Decomposes work into possible projects | Accepts responsibility for undertaking them |
The machine supplies persistence, breadth, association, and repeated critique. The person supplies motive, situated judgement, and responsibility. This makes memsearch a concrete case of the many-eyes argument in the generative AI curriculum page: the value lies in several different perspectives meeting, and the responsibility for the outcome stays with a person.
Failure Modes and Epistemic Risks Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A case study is only as useful as the risks it names. Several are already visible in the design.
- Representation failure. Embeddings can make superficially similar passages look related while separating concepts that are connected but described in different vocabulary.
- Memory pollution. Synthesis reports are mined back into the store. A later search can return a generated report beside original evidence, with no structural difference between them. Over repeated cycles, speculation could become hard to tell apart from source material.
- Self-reinforcing synthesis. If a generated idea is re-ingested, a later synthesis may treat the earlier speculation as evidence that a relation is important. The loop that makes the system cumulative is the same loop that could amplify a weak idea.
- Verification theatre. A verification stage can look rigorous while checking only details that are easy to search for, or while accepting sources that loosely resemble a claim.
- Goal drift. A system may optimise for ideas that look interesting rather than ideas that are useful, feasible, or true. Nothing in the pipeline currently measures usefulness.
- Salience bias. Heavily represented projects can dominate the embedding space, while small, recent, unusual, or poorly documented projects are hard to retrieve.
- Privacy and boundary failure. The tool is local-first, but verification sends queries derived from private material to an external service, so some exposure remains.
- Premature project formation. Consolidation and decomposition can give a speculative idea the look of maturity by wrapping it in a project name, work packages, and interfaces.
These risks do not invalidate the project. They define the controls and evaluation questions that the next version needs. The brittleness described on the agents-that-fail page applies here too. The store holds only what has been mined, so a question far from that material may be answered with confidence and little basis, which is an inference from the architecture rather than a measured result. The same applies to output that misrepresents its own status when a generated report sits beside original evidence with nothing marking the difference.
What Would Need to Be Evaluated Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
At present the project has no evaluation against a standard that would make its output a sound basis for research conclusions. A useful evaluation would ask whether the candidate ideas that survive verification are judged useful by people who know the domain, whether verification finds genuine errors in generated claims, whether generated material can be told apart from source material by someone reading the store, and whether the system's suggestions change what its user actually does. Usage counts alone cannot answer those questions.
Questions the Project Leaves Open Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The project is unfinished, and its open questions are a productive part of the case.
- Should generated material be stored apart from primary sources?
- Should every derived idea carry a complete record of its provenance?
- How should confidence change after verification, and in which direction?
- Can the system be made to seek disconfirming evidence actively, rather than only checking claims it has already made?
- When should refinement stop?
- Who or what decides which idea receives attention?
- Can the system recognise duplication with an established field described in unfamiliar language?
- Should actions that change persistent state require explicit human approval?
- How can usefulness be evaluated without rewarding novelty alone?
- Can the system explain why two clusters were treated as distant or as bridgeable?
- What should happen when local knowledge and external evidence conflict?
- How should forgetting, expiry, and correction work for generated material?
These questions turn an incomplete implementation into a design enquiry, and they match the questions the auditable-agents page raises about where checkpoints belong.
Closing FoundationalKnowledge that endures for decades — core principles
Memsearch did not start out as an attempt to implement a cognitive architecture. It started with the practical problem of finding and connecting material scattered across one person's work. The need to take in experience, keep several representations of it, retrieve relevant episodes, abstract patterns, imagine connections, test them, revise them, and turn some of them into future action has nonetheless produced many of the same components. That convergence does not show that the system thinks, and it does not show that every search tool with a language model is an agent. It does suggest that cognitive architectures sometimes emerge from confronting the functional problems that cognition solves, rather than from imitating cognition directly.
Related Topics
- Sense, Model, Think, Predict, Imagine, Act: The Cognitive Cycle Behind Every Agent — the shared diagram this case study maps the system onto.
- Agent Archetypes: Five Tiers on One Diagram — the tiers that memsearch's functions cut across.
- When Agents Fail: Brittleness, Misaligned Incentives, and Deception — the failure modes the risk section applies to a concrete tool.
- Designing Auditable, Robust Agentic Systems — the design responses, including the separation of verification from generation.