Quorumz Research

Atomic Alignment: results, contributions and limits

Eitan Lavi · Quorumz Research

Transcript · Synthetic English narration · Based on the Quorumz whitepaper and September 2026 research manuscripts

Watch the film

Atomic Alignment

CONSOLIDATED RESEARCH

Our updated framework separates two learning responsibilities. A2RL trains the personal core around its owner. Recursive cooperation learning trains each iCore's own choices about peer help. One prompt can create a sequence of referrals, a live quorum, or a combination of both. We distinguish these proposed mechanisms, their conditional mathematical guarantees, and the older synthetic measurements.

Source: Abstract

What is our contribution?

ORIGINAL WORK AND PRIOR METHODS

Expert models, learned gates, decentralized routing, and reinforcement learning all have precedents. Our work brings continuing owner feedback, complete personal agents, local recursive decisions, and independently enforced authority into one contract. We also construct a finite learning example with a proved advantage. We have not established a literature-first claim.

Source: 2.1 What conventional experts mean

A2RL learns the personal core

A2RL: A LEARNING AND AUTHORITY CONTRACT

A2RL now refers specifically to personal-core learning. It covers owner-endorsed preferences and writing behavior, task completion, and acquired skills. Memory, demonstrations, corrections and verified outcomes stay different data types. A document is not automatically a reward. Character alignment means specified behavior the owner endorses, rather than proof that the model has recovered a unique inner personality. A separate cooperation policy learns how to obtain useful help.

Source: 2 The personal core

What happens in a nightly review?

PROPOSED WORKFLOW / CONDITIONAL RELEASE BOUND

During a proposed nightly review, admit permitted experience and separate memory from reward evidence. Freeze the interpretation of feedback. Train a personal-core candidate against fixed cooperation, then a cooperation candidate against fixed cores and peers. Evaluate the actual pair independently before release. Two components can each improve with the old partner and fail when combined. A joint check is therefore essential, and the nightly schedule itself provides no guarantee.

Source: 8 Nightly consolidation

Every iCore decides its next step

LOCAL RECURSIVE COOPERATION

There is no required global semantic directory. Each iCore uses its own peer references, competence evidence, and authorized referrals. It can answer, forward, split a task, combine replies, or stop. Familiar contacts can shorten a search, but social closeness establishes neither skill nor permission. Broad questions follow the same local process.

Source: 4.1 Forward split join and return; 4.5 When natural quorums reduce search cost

Learning which connections help

ESTABLISHED POLICY GRADIENT APPLIED LOCALLY

Cooperation learning scores the completed task, including later peer contributions and costs. A local call may cause many further calls, so a quick first reply is not enough to judge it. Under a frozen finite execution, each personal action and each cooperation decision contributes its own policy-gradient term. Repeated calls to one iCore count again. This uses established coagent reasoning, and does not prove convergence when many neural agents change simultaneously.

Source: 4 Learning the two policies

Sequential referrals and live quorums

RECURSIVE EXECUTION

A sequential referral asks one peer, which can ask another and return a result. A live quorum forks separately budgeted subgoals to several peers and joins their replies. Each member may recursively form another collaboration. The same prompt can combine both modes. Local task models do the reasoning and synthesis; the cooperation policies decide how to organize that work. Every round needs a coverage rule, permission and a stopping condition.

Source: 3 Connectivity

Calls branch. Replies come back.

CAUSAL EVENT GRAPH: R1

The persistent peer graph may contain cycles. A particular execution is described by separate send, receive, join, and finish events, ordered by cause. Replies return along recorded calls and may be combined on the way. This event graph is acyclic. Drawing calls and replies as opposite arrows between whole-agent nodes creates cycles; separating their events avoids that.

Source: 4.1 Forward split join and return

A cascade cannot invent more budget

CONDITIONAL ENFORCEMENT RESULT: R2

A caller spends part of its grant, keeps part, and delegates the rest. Trusted local monitors must prevent duplicated or rolled-back spending, and retries count too. Adding those local constraints bounds total charged resource use without a global ledger. Signatures alone do not enforce this. Deadlines and availability are separate requirements.

Source: 4.2 Conserving a query budget through local grants

What the earlier tests actually showed

ARCHIVED MEASUREMENT / ORACLE DIAGNOSTIC

The best specialist scored fifty-five point six percent, the three-specialist merge sixty-six point seven, and the best merge sixty-seven point eight. Perfect selection reached ninety-three point four percent, an oracle ceiling rather than a working router. The roughly twenty-nine percent relative merge loss against the oracle ceiling was one three-pack comparison, not a universal tax per pack.

Source: Existing evidence and explanatory limits

Private expertise can combine

E4 / COMPLETED SYNTHETIC MEASUREMENT

In the completed E4 study, three privately trained iCores answered forty-five point one four percent of joint questions, versus eighteen point seven five percent for the selected standalone. Removing a necessary owner sharply reduced complete reasoning traces. This is evidence of complementary learned knowledge in a controlled synthetic setting, using compatible branches and one decoder. It is not a frontier or pooled-training comparison.

Source: Frozen evaluation and results

Relevant growth must preserve competence

E4 / NEGATIVE SCALING RESULT

On the exact same original questions, adding irrelevant members reduced joint-answer accuracy from forty-five point one four percent at three members to twenty point eight three percent at twenty. The twenty-member primary result failed its registered threshold. A router-only intervention partly repaired performance, but did not solve scaling. More registered iCores are not sufficient evidence of greater intelligence.

Source: Holding the original cases fixed

Compose in the right place

MEASURED LIMITS / CONDITIONAL DESIGN RULE

Facts can remain in permitted retrieval stores. Conflicting behaviors can stay in separate specialists. Weight merging needs evidence that each task survives the actual update. Our conflicting-cipher and failed linear-graft results constrain the theory. A task-loss bound supports conditional merge decisions; it does not promise that every useful skill can be merged.

Source: 6.2 Three different composition operators; Existing evidence and explanatory limits

A proved advantage, with its price visible

PROJECT-SPECIFIC FINITE CONSTRUCTION: A2.2

In our constructed task, every owner supplies twenty-four informative outcomes. A central learner receives at most one per owner on average, with no other owner-specific information. Owner learning with separately trained routes yields an expected-reward gap above zero point zero six one three, before execution penalties. This is an information-access advantage, not an equal-data efficiency proof.

Source: 5.2 A finite owner learning and routing separation

Two distinct sources of answer quality

Q1 / CONDITIONAL THEOREM

Our quality bound separates the shortfall of the personal cores from errors in cooperation. Cooperation can lose value by missing a useful peer, estimating its contribution incorrectly, or choosing a weaker option. The bound compares complete task interfaces, not unrelated standalone accuracy. A network beats a specified comparator only when its available competence advantage exceeds these losses. For subjective writing, owner ratings still need a justified interpretation.

Source: 5 A2RL quality

What does fractal mean here?

RECURSIVE INTERFACE RESULT: C1

A single iCore and a group can expose the same task interface. Replacing a group with its exact complete boundary behavior preserves composition. This gives recursion a precise meaning. It does not prove that nested voting equals flat voting, that compressed summaries preserve every dependency, or that a brain-like architecture emerges.

Source: 6.1 Quorums share an interface rather than identical internal weights

A part can recover authorized coverage

SPECIFIED HOLDER MODEL / EQUATION 6

Our reconstruction result is narrower than the word holographic suggests. For each authorized item, a specified decoder succeeds exactly when a designated holder is available within the permitted access budget; otherwise it abstains. More copies improve expected coverage under independent availability, while costing storage and potentially increasing exposure. This does not prove arbitrary reconstruction of private knowledge.

Source: 6.3 A falsifiable meaning of holographic recovery

Atomic control needs a trusted boundary

COMPANION AUTHORITY PROOFS

Every owner's protected controller checks access and actions independently of learned weights. Under complete mediation and protected state, later local actions can be blocked after revocation. This proves a control property under assumptions. Owner approval is not proof of factual correctness, faithful rewards, or alignment with everyone affected by an action.

Source: 9 Authority and scientific limits; Consent as a compositional authority invariant; Atomic local revocation and bounded remote containment

Some privacy limits are unavoidable

INFORMATION AND REVOCATION LIMITS

If two private situations require different answers but become the same shared question, with no other permitted information to distinguish them, a peer cannot always answer correctly. Removed information cannot be recovered by better routing alone. Models and relationship graphs can also leak information. Remote permissions need bounded expiry, and a recipient's existing copies cannot be erased by a later stop command.

Source: 9 Authority and scientific limits; Query minimization and its sharp limits; Information-flow containment and adaptive compromise

Many available. Few activated.

E1 AND E2 / CONDITIONAL BOUNDS

Suppose an entry invites a small first quorum. If the expected number of later invitations contracts at every generation under the actual task history, expected total activations have a finite bound. Failed calls and retries count too. Energy savings require this bounded work, communication, and the allocated cost of training to stay below the comparator. Low fanout among idle registered nodes is not enough. Hard budgets still provide the deterministic limit.

Source: 6 How many

Faster replies do not mean less work

E3 / CONDITIONAL LATENCY RESULT

When independent required subtasks can run together without resource contention, a quorum waits for the slowest complete reply. Serial execution waits for the sum of those reply times. That can reduce latency, but the same work still consumes resources. A sequential search that can stop early is a different comparison. Shared devices, congested links or one mandatory slow member can remove the speed advantage.

Source: 7.1 Sequential

Scale requires a complete cost account

CONDITIONAL THRESHOLDS / CENTRAL EMULATION

Millions of available experts do not imply millions of experts should answer each question. Exact fresh consultation may require contacting everyone. Energy improves only when avoided costs exceed local learning, discovery, communication, and other added costs. An equally informed and provisioned central implementation can emulate the network. Universal decentralized superiority is therefore ruled out.

Source: 7 Scale and efficiency thresholds

Two learners. One accountable release.

PROOFS, MEASUREMENTS AND OPEN WORK

The framework now distinguishes personal A2RL from recursive cooperation learning. Its conditional results separate core quality, routing losses, activated work and response time. The proposed implementation records permitted experience, versions, grants, replies and complete outcomes so those conditions can be checked. Existing scaling failures remain evidence. The next scientific step is to establish these contracts on declared workloads and compare the complete costs with equally informed alternatives.

Source: 9 Implementation contract