Skip to content

fix(web-forms#883): allow users to recover from xpath errors - #1908

Open
garethbowen wants to merge 41 commits into
getodk:masterfrom
garethbowen:883-node-value-result-type
Open

garethbowen wants to merge 41 commits into
getodk:masterfrom
garethbowen:883-node-value-result-type

Conversation

@garethbowen

Copy link
Copy Markdown
Contributor

Closes getodk/web-forms#883

What has been done to verify that this works as intended?

Manual testing, CI.

Why is this the best possible solution? Were any other approaches considered?

Discussed on the issue.

How does this change impact users? Describe intentional behavior changes from code updates. What are the regression risks?

Error messages are now shown without interrupting flow, and inline if possible.

Does this change require updates to user documentation? If so, please file an issue here and include the link below.

No.

@changeset-bot

changeset-bot Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: a826b71

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 3 packages
Name Type
@getodk/xforms-engine Minor
@getodk/web-forms Minor
@getodk/xpath Minor

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

fields: references.join(', '),
count: references.length
}) }}
</li>

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The most common case is a single error violation, potentially impacting multiple fields, so I've optimised the UX for that case.

@@ -91,14 +91,23 @@ describe('#format-date()', () => {

describe('invalid dates', () => {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Now some of these throw and some return null... It may still not be right but it's closer.

@garethbowen

Copy link
Copy Markdown
Contributor Author

@latin-panda This is now ready for review, but given the size the change ended up being I'm not sure we want to merge it in 1.1.1. I'm happy to leave this for 1.2.0 if that feels safer to you.

@latin-panda latin-panda left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I haven't finished reviewing or testing, but I wanted to share my feedback so far. I'll continue tomorrow

Comment thread packages/xforms-engine/src/instance/abstract/DescendantNode.ts Outdated
const chunks: TextChunk[] = [];
const mediaSources: MediaSources = {};
const chunkExpressions = getChunkExpressions(context, definition);
context.setError(null);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let me see if I'm getting this correctly:

Each node/field has one error "slot". And many things write to that slot (the field's calculate, relevant, label, constraint). Each of them sets the slot to null when it succeeds.
So the label succeeds and sets it to null. But if the calculate failed before then, the error is gone. Does it make sense? 🤔
Maybe it's better to have a map of actors (calculate, constraint, etc.) and the corresponding error.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done, thanks for spotting this!

My solution feels clunky so if you have a better way to do it I'd love to hear it, but this does work.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you! The map is better, but each place still has to set and clear its own key. I haven't checked all cases but I found that createTextRange L111 sets the label error and clears it on the line below, so a broken itext shows nothing, and label, hint and messages share the same key, so one can hide the other's error.

I was having a deeper look, and each expression is already a memo that returns a Result, and result.error.message has the text we need. So we could have one memo per node that reads all the results and returns the ones that failed. Since it's reactive, it re-runs when something changes and then a fixed expression returns success, the error is removed automatically.
The createErrorValidation can read that memo instead of getError(), and the violations list stays the same.

The only thing I see so far with that approach is the calculate, setvalue and the text chunks don't return a Result yet, so they'd need that.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤔

It's definitely a more reactive approach, and solves all the issues about keeping the error state correct. If I understand correctly, the InstanceNode would have an abstract errors accessor that child nodes would override with a memo which reactively updates based on the computation Results that it knows about. For example, the ValueNode would check for errors on the value Result, and the RepeatRangeControlled could check for errors on the computeCount Result. Each Node's accessor would probably also call the accessor on their parent node to get a combined list of errors. Is that what you were thinking?

It's another significant change so I wanted to confirm I understood what you meant before spending too much time building it out.

@latin-panda latin-panda Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, correct! Just one thing about call the accessor on their parent, if you mean the parent class (super.errors() + the node's own results), yes. But if you mean the parent node, the createAggregatedViolations already walks the children up to the root, so pulling from the parent would list the same error once per descendant, right?

Thinking about setvalue from my last comment, the actions run on events, so there is no Result to read. They would need to store the error (or null) of their last run on the destination node, and the node's errors() reads that 🤔

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if you mean the parent class (super.errors() + the node's own results), yes

Yes this is what I meant.

Thinking about setvalue from my last comment, the actions run on events, so there is no Result to read.

Yes I think there will need to be some calculations that store errors outside of the Result. I'll see what I can come up with today.

Comment thread packages/web-forms/src/components/OdkWebForm.vue Outdated
@@ -339,7 +343,11 @@ const showValidationError = computed(() => {
if (errorBannerDismissed.value) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I haven't tested it, but I think errorBannerDismissed should be set to false here so that new evaluations with errors display the banner.

@garethbowen garethbowen Sep 23, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't currently do that for other violations until you hit the next page or submit button, right? I feel like it would just be annoying if you have explicitly dismissed it for it to pop up again, even if now the message is different.

@latin-panda latin-panda left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still digging into the details :) Meanwhile, I wanted to share some feedback, and I replied to one of the threads with an idea for catching errors from Result more reactively. Let me know what you think!


case 'upload':
// leaf node
return violationReference(child);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In L38 collects attributes for the parent, right? but a leaf question never becomes context, so its attributes are skipped (the submit is not blocked on error)

Suggested change
return violationReference(child);
return [...violationReference(child), ...child.getAttributes().flatMap(violationReference)];

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

True! I've fixed this by moving the attributes violations into the violationReference function, so when it's getting violations for the node it also checks the attributes.

readonly evaluator: EngineXPathEvaluator;
readonly contextReference: Accessor<string>;
readonly getActiveLanguage: Accessor<ActiveLanguage>;
readonly errorState: Accessor<ErrorState>;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I couldn't find anything that reads this error state. When an item label is broken (<label ref="badFn()"/>) the error seems lost?. I wonder if it could bubble up to the control's state.

message: error,
} as const;
}
return null;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't block the form if the question is relevant=false, right?
I think it should return null in that case
if !context.isRelevant() { returns null

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm.. In this case I'm not sure. One example is what if the expression for relevant is the thing with the error? Ultimately if this field is in an error state we can't be confident if it is relevant or not. To be on the safe side I think it's fair to block any form with an unresolved error in it.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good! If the question isn't visible and the relevant value is false, could this be used to improve the error message?
"the hidden field has invalid ....."
Or the question name/ref.

@garethbowen
garethbowen force-pushed the 883-node-value-result-type branch from 3307a21 to 55326b8 Compare September 23, 2026 22:55
@garethbowen
garethbowen force-pushed the 883-node-value-result-type branch from 8818fdd to 4cf9e96 Compare September 28, 2026 21:09
@garethbowen

Copy link
Copy Markdown
Contributor Author

@latin-panda This is one of those changes that keeps going, and I'd appreciate a gut check to see if it's worth continuing. I've implemented your suggestion to make the error handling more reactive, rather than storing it in a map, and it's now reached a stable point in this PR. I have yet to handle errors for node options, labels, and hints. These are all significant additional changes so I paused to do some re-evaluation if this is the right approach.

Firstly the code is significantly more complicated after this. Just about every file in the engine now needs to have additional branches to handle errors that may be reactively returned on any update.

Secondly the engine is somewhat slower. There may be some optimisation missing but as it currently stands, the two child-vaccination smoketests run in 7.5s and 22s, up from 6.1s and 18s. So that's about a 20% slow down. It does make sense that it's somewhat slower because every change now needs to be checked to see if it's a failure, and the validation state checks many more potential sources of error.

On the other hand, a definite benefit is it will give us a clean way to handle error states that occur within the Engine itself which are currently still throwing.

Because this should be a very rare occurrence I no longer believe it's worth the performance and complexity cost.

What do you think?

If you agree, then we have a couple of options, in addition to those I outlined in the issue.

  1. Leave the errors being thrown up to the client (status quo, close this PR). We could investigation doing some pre-render validation, and improving error messages where possible (ie: wrapping the error in another error which includes additional context like the nodeset).
  2. Still catching the error at the engine boundary, but instead of returning a Result, set an error on the root node and return a default value. This error could be shown in the top banner but still allow for users to keep typing. The error wouldn't be shown inline but in many cases it cannot be anyway, for example, non-body elements.

@latin-panda

Copy link
Copy Markdown
Collaborator

@garethbowen, thanks for all the work on this, and for being open to trying all these ideas!

I was curious about where the slowdown comes from, so I dived into the code and ended up trying an idea on top of your branch. I didn't want to step on your code, so I put it in a separate draft PR for you to look at, no pressure to use it (the PR is not complete, it needs more test coverage and more polishing).

I found that every evaluation returns a new Success object, so the memo always thinks the value changed and everything downstream recalculates (about 4x more evaluations than master, from what I could track in my tests). The draft PR keeps your client types and the UI, and changes how the error is kept, in the expression memo, next to the value, so Result doesn't need to travel through the signatures. It also fixes a memo leak in createTextRange that is already on master. To give an idea, these are the two child vaccination smoke tests (jsdom on my machine):

  • test 1
    • master: 2s, this branch: 3s, draft PR: 2s
  • test 2
    • master: 5.5s, this branch: 6.5s, draft PR: 5s

What do you think?
If you agree, then we have a couple of options, in addition to those I outlined in the issue.

I agree it's not worth it with the complexity and the slowdown, but I think the draft PR removes both. It's actually close to your option 2. The difference is that the error is kept on the node and not on the root, so it clears by itself when the input is fixed. It also covers labels, hints and node options.

If you prefer to keep your approach, there is also a small change that helps a lot. The createMemo takes an equals option to decide if the result changed. The default is ===, and a new Success object is never === to the old one, so the memo always notifies. Passing a function that compares value and error.message stops that, and it runs faster. However, this still leaves the extra memos on every node and the cost of the walk that collects violations, which the draft PR reduces.

return this.getValidationViolation();
}

// Attributes never change once the node is built, so they are read without tracking.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In which case, shouldn't the getAttributes function also be untracked? Should attributes just be an array removing the need for the createAttributeState call at all? This is probably outside the scope of this change but could be a nice performance and simplicity win.

// prettier-ignore
type ComputedExpression<Type extends DependentExpressionResultType> = Accessor<
EvaluatedExpression<Type>
export type ComputedExpression<Type extends DependentExpressionResultType> = Accessor<

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the main difference between your draft pr and this... In the draft PR the computed expression had an error property with a function. I found this difficult to read, and meant there were two possible ways errors were being tracked, either through the error() fn or through the registerExpressionError

My latest commit uses the registerExpressionError exclusively so the ComputedExpression is unchanged from master.

@garethbowen

Copy link
Copy Markdown
Contributor Author

@latin-panda Thanks! I merged your draft PR and then made a few changes...

  1. I added back the Result for the engine evaluator. I'm more comfortable with this because then the caller is prompted to check for errors, rather than having to remember to try/catch. I didn't take it as far as with my previous version so it didn't cascade throughout the engine.
  2. It was slightly awkward having two different ways errors could were being reported so I standardised on using the register function (comment inline).
  3. I added a range more tests for different places expressions are executed.

This passes the tests, and executes the smoke test in about the same time as your draft PR.

What do you think?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Allow users to recover from xpath errors

2 participants