Repository navigation
If default sending relay is down/unreachable, every message sent delays #8762
Description
Activity
This problem appeared somewhat quicker than expected, maybe we can agree already on a way to fix it for now? I see three options:
- Always choosing the one that was used for sending last
- Keeping statistics of relay speed
- Starting with the relay that downloaded a message most recently: It was the fastest relay that successfully received this message, so it is probably also the fastest relay that works
Solution 1. is straightforward, requires a little bit of extra state; 2. requires further discussions, so it's nothing for now; 3. would require no additional state (we already have
last_rcvd_timestamp), but IIRC @link2xt was against using receiving behavior for choosing a sending relay since sending and receiving behavior of a server is sometimes different?(BTW, good that we didn't choose to try sending relays in random order; otherwise, it would have been way harder for the users to debug the problem, and while some messages would have gone out quickly, it would have affected also users who have this relay, but not as the first one)
Always choosing the one that was used for sending last
I think the simplest solution is to remember the timestamp of last successful sent message for each SMTP transport and sort the transports by this timestamp rather than by rowid. For transports with the same timestamp we can still sort them by rowid, it should not really matter in most cases how they are sorted.
There is also a somewhat related topic on the forum, just for reference: https://support.delta.chat/t/2-62-broke-mixed-use-of-regular-email-and-chatmail-relays-for-one-identity/5838
I don't think we want to support the case of sending unencrypted messages mixed with chatmail relays, but there are other messages in the topic and https://support.delta.chat/t/2-62-broke-mixed-use-of-regular-email-and-chatmail-relays-for-one-identity/5838/5 describes a similar problem ("Also, after a few tests (at least in my case), DC keeps trying with the previously set sending relay (set with previous DC version). In case of error the message is not sent and nothing can be done about it. The automatic switching is not happening"), but maybe it just means that the message is sent with the last relay (chatmail) and the message is not encrypted so it is rejected with an error.I created #8771 to always use the transport that was most recently used. It is good enough to close the issue.
For something fancier, we have three layers that can be reordered:
- Transports
- Hosts+ports (what is called "ConnectionCandidate")
- IP addresses after DNS resolution if it is a domain name
Last layer (DNS resolution, TLS, SNI etc.) i would rather leave self-contained, but the first two layers with some refactoring can be merged and then we should be able to interleave "connection candidates" from multiple transports. E.g. try port 993 from the first transport, then port 143 from another transport etc. Otherwise if we remember that some transport was successfully used, when we try it we have to try all "connection candidates" of this transport even if some port always times out because relay operator firewalled some port away. This will still need another table because existing
connection_historyrecords successful connection rather than successfully sending a message, so not good enough (worse than #8771) for ordering SMTP candidates, but tables are cheap to add anyway.I had curiosity of how this process worked and one question I have is if DC choosed the shortest path, like if I want to send a message to someone I have the same relay configured if it would use that relay to send it. Does it have any sense or it's just a human logic that technically doesn't makes sense?
And another question, if DC checks the connectivity with a relay and knows the status of each one, would be possible to avoid to send messages from the ones that are down?
Is simplicity in this decision making a priority? Because It would be interesting a more elaborate process like:
1- Is the relay up?
2- Do I have configured the relay of the receiver? (if that makes sense, as mentioned before)
3- Is enough space left in this relay account?
4- Is the relay able to send this message? (I want to configure my relay to send bigger files than the default setup)for example...
Reported from an operator: they had their relay down for some time, and users who had it as their first relay then got 1-2 minutes delays when they sent a message. It's not a surprise because smtp is currently always tried in the same fixed order, every time. But it's a bug in the user experience, and somewhat of a downgrade from before 2.62 where users could influence what is the sending relay.