TechDex Research documents technical work involving governed AI, cognitive architecture, epistemic behavior, persistent operational identity, and sustained conversational intelligence. Publications are evidence-led, versioned, and explicit about their evaluation status, limitations, corrections, and relationship to peer review.
Public working preview - updated for the September 12, 2026 session. This research remains in development and has not been peer reviewed or formally published as a completed paper. The preview will be maintained as implementation and evaluation progress; public availability does not imply completion or independent validation. The intended destination is a research manuscript submitted for external peer review. No acceptance or submission readiness is claimed at this stage.
Subsequent evaluation now includes the Framework 1.1.11 acceptance review: 66 turns in total, with 60 turns across five main conversations and six additional focused probe and browser turns. Human adjudication graded the five main conversations as three passes, one bounded pass, and one partial. Observed improvements include compound-request preservation, disagreement that retains the requested task, correction continuity, and explicit closure. Partial-result handling, a replayed choice discrepancy, correction fallback, and qualified source identity retain defects or verification gaps.
At the September 11 evidence stage, Framework 1.1.12 had 32 local regression suites passing and live acceptance was pending. Subsequent September 12 work separated natural article lookup from targeted research edge cases. Focused live checks through 1.1.15 observed improved handling of declarative follow-up questions, while retained-field reporting, source selection, and source projection on continuation and stopping still showed defects. Those focused checks do not establish extended-research reliability.
Framework 1.1.16 is now published for installation. Its repairs address single-source operation scope, current-task cessation, and article-result eligibility. All 36 local Framework regression suites pass, including 30 new checks. A subsequent fifteen-turn focused live check on TechDex found clean explicit stopping and useful conversations that recovered through clarification and correction. Retained-link reporting and unrelated-source presentation still showed defects; the targeted research objective remained unmet. These short tests do not establish extended-research reliability. Local contract validation is not proof of improved customer conversations, and no higher conversational rating or independent validation is claimed.
The evaluation now distinguishes conversation usefulness, task outcome, retained defects, and successful repair. Earlier saved conversations were regraded under that explicit standard; a regrade changes interpretation, not historical software behavior. These later evaluations use different instruments from the fixed suite below and their scores are not combined.
Later focused checks on Framework 1.1.17 verified saved partial-result identities and Ledger progress updates, while exposing an undertaking that remained open after a subject change and stopping. Framework 1.1.18 addresses that lifecycle boundary with stable undertaking references and verified Ledger receipts before finishing is confirmed. Its 38 local regression suites pass. Seven subsequent focused live turns verified stopping across subject changes, persisted finishing receipts, and a new task completing in the same conversation without reopening the stopped work. Source-selection and requested-date interpretation defects remain. These stages remain draft research evidence, not independent validation or a claim that extended research reliability has been established.
Framework 1.1.19 addresses a source-reference admission defect: a follow-up about the same article could admit a different article when no antecedent had been established. A resolved source identity now constrains admission and reuse; unresolved references cannot acquire a substitute through word overlap. All 39 local regression suites pass, including 28 new checks. Live acceptance remains pending. This does not establish universal source relevance or a higher extended-research rating, and the separate requested-date interpretation defect remains open. These are draft, internally tested findings.
Framework 1.1.20 adds event-specific date interpretation: asking when an article was checked or reviewed must preserve that event, rather than request its author or substitute its publication date. Non-temporal comparisons and causal language no longer acquire date constraints from a preposition alone. All 40 local regression suites pass, including 29 new checks. Live acceptance remains pending; these findings do not establish broader research reliability.
A subsequent 13-turn focused live check (including two onboarding turns) verified the saved source-reference repair and event-date distinction on TechDex, including retained article links, useful continuation and committed stopping. Fresh Live Minder testing still exposed an unrelated source on an initial search, despite the response correctly recognizing the mismatch. These results support the bounded repairs; initial-search relevance remains open. They are not an extended-research evaluation or independent validation.
Framework 1.1.21 extends the language representation to distinguish an explicitly attached source subject from source location and response-format instructions. MiniBrain carries that subject into evidence admission. All 41 local suites pass, including 33 new checks that reproduce the Live Minder mismatch and vary wording and subjects independently. Live acceptance remains pending. Relevance screening still uses lexical coverage within the resolved subject; this is not a claim of general semantic entailment or complete grammatical coverage.
Nine subsequent live turns support the 1.1.21 admission repair: the original unrelated source is excluded, while relevant article retrieval and TechDex continuity remain available. A compound positive control exposed a separate operation-scope defect: an explanation's ordering word changed retrieval chronology. Removing that clause restored relevant retrieval. The compound conversation remains partial; this is not a general retrieval success claim.
A subsequent local candidate carries operation identity and dependencies through retrieval, source qualification, body delivery and completion accounting. Forty-three local regression suites pass, including 49 new contract and database-pipeline checks. The saved compound request and its continuation are covered by deterministic fixtures; these results are not a live conversational evaluation. The repair is published as Framework 1.1.22; live acceptance is pending installation. Earlier grades remain unchanged, and this research remains a working draft, not peer reviewed.
Subsequent 1.1.22 live acceptance covered thirteen turns across four short conversations. Human review found one pass, one bounded pass and two failures. The original compound retrieval improved, but next-turn source recovery failed and fresh wording exposed an unrelated projected source. The offline anchor fixture did not reproduce real response-to-session persistence. Consolidated acceptance remains unsuccessful; extended-research ratings are unchanged.
Framework 1.1.23 represents the purpose of indefinite source requests and nested activity clauses before source qualification. Local tests exclude the saved unrelated source before evidence registration while retaining the relevant article. All 44 suites pass, including 312 wording-and-subject combinations repeated with five candidate-order seeds and successive admission checks. These are deterministic tests, not live conversations; the repair has been published for installation but has not received live acceptance; source-authority persistence remains separate.
Focused live acceptance of 1.1.23 covered twelve turns across five threads. The saved unrelated-source failure now returns only the correct article, and a second-site positive control succeeds. Broader delivery remains incomplete: a nested request retrieves the correct article and reaches synthesis with accepted evidence, but the response boundary reports the source missing. This distinguishes the observed admission improvement from unresolved evidence handoff and response accounting; no overall maturity claim follows.
A subsequent repair published as Framework 1.1.24 traces that failure to a projection check that contradicted MiniBrain's authorized operation. Projection now follows the operation binding, while acquisition and delivery are accounted for separately. Connected checks also cover Ledger operation identity and session preservation of execution receipts. All 45 local suites pass; the handoff suite includes 106 checks, of which 33 belong to the reused retrieval fixture. These results are local integration evidence, not live acceptance or a claim of sustained research reliability.
Subsequent live acceptance of 1.1.24 covered twenty diagnostic turns across three conversations, including onboarding and stopping turns. The exact nested-request replay succeeded and a user correction persisted in another conversation. Broader acceptance did not pass: unrelated source projection recurred, and ordinary writing and explanation requests were incorrectly blocked by a source prerequisite. A fresh conversation reproduced the writing failure. Runtime traces show reasoning authorization and completion accounting disagreeing about required evidence. These live findings remain distinct from the passing local regression suites; no overall maturity claim follows.
A subsequent role-alignment review examined all four installations' public environments, eight domain and identity exchanges, authenticated workspaces, active reporting connections and an existing saved business report. The evaluation target is useful, governed conversation within each installation's actual business or educational role. The earlier standalone reminder-writing probe remains an internal diagnostic observation, but is not a primary product-acceptance requirement. Source relevance and conversational continuity remain pertinent findings. This clarification preserves the earlier evidence and assigns no new performance score.
The common evaluation standard is the Framework's purpose: coherent, helpful conversation that makes information useful. Installation roles provide context rather than separate grading standards. A subsequent code and saved-evidence review identifies inconsistent evidence requirements across coordination phases and incomplete separation of conversation and undertaking lifecycle. It also identifies an automatic task-reopening rule in Ledger recording that requires explicit coordinator ownership. Useful explanations, persistent corrections and verified stopping remain positive evidence; exact attribution of every saved-state anomaly is still pending.
Framework 1.1.25 reconciles optional and required source evidence across coordination, connects associative follow-ups to established governed sources, and distinguishes stopping work from continuing conversation. MiniBrain adjudicates Ledger transitions; persistence checks that the underlying state has not changed. Corrections retain scoped history without automatically reopening unrelated completed work.
Local validation passed 46 regression suites, including 42 new purpose checks and four additional checks through the actual response preflight. Installer and updater checks passed. These are deterministic local results, not live conversation counts. The release is published; client installation and live acceptance remain pending. No conversational ratings changed. This remains a working draft, not peer reviewed.
A subsequent ten-turn conversation on one verified installation showed useful ordinary explanation, correction, subject return and visible stopping of research while conversation continued. Source continuity did not pass: unrelated articles remained in a follow-up, the corrected response and a stop acknowledgment. Local regression success therefore does not establish complete live conversational repair. Ledger commits were not independently inspected in this run, and no numerical ratings changed. These findings remain preliminary and not peer reviewed.
Follow-up tracing places one unrelated source before answer synthesis: it received evidence eligibility and an operation receipt. The preceding successful source request had not established the conversational anchor required by the follow-up path. A local mechanism probe reproduces that missing establishment and incorrect admission. The prior fixtures assumed the anchor already existed. Full live-state reproduction remains incomplete; these observations support a continuity-to-admission defect rather than a purely visual citation defect.
MiniBrain now derives conversational source identity from accepted operation evidence and final source projection. Qualified source metadata survives anchor storage so the existing governance boundary can authorize continuation. Composed local tests exercise retrieval, response commit, storage and the next turn's source reader and admission, rather than supplying an already-authorized anchor.
All 47 regression suites passed, including 36 new checks. Repeated local follow-up admission retained the established source while rejecting unrelated candidates. The release is published; live acceptance after client installation remains pending. These deterministic results do not establish unrestricted source continuity or a new conversational rating. This is a working draft, not peer reviewed.
Working draft; not peer reviewed. In one 10-turn conversation on an owner-confirmed updated installation, the original source-continuity replay succeeded: the opening source was established and used on the follow-up. A fresh question about research funding was instead treated as an author-information request, and immediate clarification did not repair it. Human assessment: partial conversation and partial objective completion, despite useful later guidance. The stop reply had no unrelated citation; durable closure remains unverified. The final new-topic response was cached and does not establish fresh source-selection performance. This was an adaptive evaluator-authored probe, not a randomized cohort or an automated-judge score. No general reliability increase is claimed.
Working draft; not peer reviewed. Local validation now distinguishes requested information from incidental mentions and explicitly rejected alternatives. Compound interpretation preserves actual field requests, while source handling consults the governed antecedent established in persisted state. Forty-eight runtime suites pass, including 49 newly added checks across request interpretation, response validation and composed source continuity, alongside installer and updater checks. This uses bounded grammatical and semantic coverage; it is not a complete language parser. Release 1.1.27 is published. Live conversational improvement remains unverified pending installation and acceptance testing.
Working draft; not peer reviewed. In one adaptive 10-turn conversation, the repaired interpretation no longer invented an author request, and explicit clarification produced a useful answer. Fresh wording also succeeded. The primary understanding objective was achieved collaboratively; requested bibliographic details remained unresolved. Full architectural acceptance remains incomplete: source binding worked on one follow-up but was released on the next, where hypothetical framing also blocked reasoning. This was not a cache hit. Visible stopping succeeded, but durable closure remains unverified. These results do not establish an extended-research rating increase.
Working draft; not peer reviewed. A new adaptive conversation completed fifteen substantive turns, with two onboarding exchanges counted separately. The discussion retained earlier distinctions, returned from a side topic, adapted to a revised user goal and separated established principles from unknowns. However, unrelated sources persisted through corrections, and a later request for the original guide returned unrelated URLs even after another correction. Human assessment: partial conversation and overall objective; educational purpose achieved collaboratively, original-guide recovery unmet. Runtime evidence showed no established article anchor throughout this thread. Precise fault attribution remains under investigation. This single conversation does not justify an increase in extended-research reliability or establish clinical factual validation.
Working draft; not peer reviewed. Analysis of the fifteen-turn run confirms that unrelated sources received operation receipts before final projection. A local reconstruction reproduces acceptance through incidental word overlap where the requested subject or historical source identity was not bound. Corrections can inadvertently reinforce the wrong matching terms. No conversational source anchor was established, and later completion accepted delivery of the wrong resource. The architectural repair proposal connects purpose interpretation, MiniBrain-governed source selection, persisted identity, correction and completion. The reconstruction is not an exact replay of private live state, and no repair or new release is claimed by this diagnostic update.
September 12, 2026 session close: Framework 1.1.27 is published and installation is owner-confirmed. The final 15-substantive-turn conversation maintained its revised purpose, returned from a side topic and continued useful explanation, while repeated recovery of the original source failed. Conversation quality and evidence continuity are therefore reported separately; this is not a claim of improved overall research reliability. The next planned architectural work concerns MiniBrain-governed establishment and preservation of the requested source object through correction, return and completion. That repair is not yet implemented. This remains a working draft, not peer reviewed; prior results and their limitations remain part of the record.
The earlier internal research record preserves the Language/Pragmatics Depth Suite: a fixed twenty-conversation x fifteen-turn evaluation, totaling 300 turns per complete run, across its original baseline and two unchanged replays. This is distinct from the earlier Source-Resumption Deep Suite, which used thirty conversations x seven turns for 210 total turns. The later Language/Pragmatics suite is a harder conversation instrument, not a replacement score for the earlier suite. The latest complete replay reached 19 of 20 strict conversation passes with one failed assertion and no transport failures. Verified article-author continuity improved from 6 of 20 baseline follow-ups to 20 of 20.
These results remain working evidence rather than peer review. The public paper and evidence package will preserve the original failures, separate machine assertions from human adjudication, disclose the tested fifteen-turn envelope, and distinguish a complete replay from later focused verification. Raw internal logs and protected implementation details will not be published.