Open model and code · Apache-2.0 code · CC BY-SA 4.0 weights
Keep decisions current as the facts change.
You say what must be true before something can be done. Messages arrive, get corrected, get withdrawn.
EASE-Delta reads each message against each requirement once, re-reads only what a change touches, and tells you
what is ready, what is blocked and why, and which single question is worth asking. It proposes; it never acts.
The reader at the heart of EASE-Delta answers supports, refutes or not enough info, and it was trained to
say not enough info when a passage is about something else. It runs here, in your browser: the text you type is not sent anywhere.
The first read downloads the model once (about 150 MB), then it is cached.
supports
refutes
not enough info
This page runs the 149M base reader, compressed to 8 bits for the browser; on 3,000 development pairs it gave the same answer as the
full model 98.1% of the time. On passages from unrelated documents, the reader gives a decisive answer 0.6–0.8% of the time; a widely used public
NLI model, in the same evaluation, 45%. Results
Watch a task stay current
A client hand-off: four requirements, two actions, twelve events, including a duplicate, a late copy of an old message, a
withdrawal and a bounced payment. Each state below is what the released 395M model produced. It is a recording; nothing was edited.
What makes it different
Learned parts read language. Exact code handles versions, precedence, logic and planning.
Re-reads only what changed
Readings are cached per (requirement, message). A change re-reads the pairs it touches: about 2–7% of the tokens of re-reading the task. Duplicates and stale copies cost nothing.
Exact, and reversible
The cached state equals a full rebuild bit for bit (0 of 195,434 values differed). A person's correction takes effect at once and can be undone exactly.
Knows when a message is irrelevant
Trained with unrelated passages as not enough info: 0.6% read as evidence, against 45% for a public NLI model.
Explains, asks, never acts
Every proposal shows the records it rests on. It names the one question worth its cost, and a READY action still waits for a person.
Measured, and pre-registered
Hypotheses and decision rules were written and hashed before any test data was read; the failures are reported too. Tasks are generated; their text is human-written.
What was measured
Result
Correct ready / blocked / needs-info decisions
91.6% standard tasks · 86.8% larger tasks
Against a dense reader of the whole task (same backbone and training)
+2.2 to +5.5 points, fewer stale decisions
Larger reader (395M) against the base reader (149M)
+2.0 and +2.2 points, close to the prediction made beforehand
Update after a change, Apple M4 Max
59 ms GPU · 0.2–0.6 s CPU
VitaminC test (claims against Wikipedia revisions)
91.5%
Conflicts that follow only from a consequence
45% recognised (a known weakness)
Use it
Python 3.11 or later. The model downloads from the Hugging Face Hub on first use.
pip install "ease-delta @ git+https://github.com/jithinsaireddy/ease-delta"
from ease import EvidenceReader, Tracker
reader = EvidenceReader.from_pretrained() # jithinpothireddy21/ease-delta
task = Tracker.define(
requirements={"approved": "Acme has approved the final design.",
"date": "Acme has confirmed the delivery date."},
actions={"send_packet": "approved and date"},
reader=reader)
task.add("mail-1", "Email from Acme: we approve the final design, please go ahead.")
task.add("mail-2", "Email from Acme: we confirm delivery on 12 March.")
print(task.status()) # ready / blocked / needs info, and why
task.update("mail-1", "Email from Acme: we withdraw our approval.")
print(task.explain("send_packet"))
from transformers import pipeline
reader = pipeline("text-classification", model="jithinpothireddy21/ease-delta-reader", top_k=None)
reader({"text": "The client has approved the final design.", # claim first
"text_pair": "Email from the client: we approve the final design."}, # evidence second
truncation=True)
# also: sentence_transformers.CrossEncoder("jithinpothireddy21/ease-delta-reader")
It has not been measured with people. Whether it saves anyone time is the next question, not a claim.
English only. Messages on the same subject that settle nothing are still misread about one time in five; say which requirement a message concerns when you know.
Requirements such as “the latest version” need two records read together and are read unsafely; name the version instead.
It reports what messages say, not whether they are true.