Conversation
Replace AsyncResolver (async DNS on separate thread) with synchronous setAddress() calls, matching the startup code path. Hostnames are now preserved in Config and Channel so periodic re-resolution and reconnect can re-resolve using the same sync mechanism used at initialization. - Add isHostname() utility to detect hostname vs IP addresses - Preserve original hostnames in Config for addressList and nameServers - Store forcedServerHostname on Channel for reconnect re-resolution - Add periodic DNS recheck timer for search destinations - Remove AsyncResolver, clientresolver.cpp/h, and dedicated DNS thread
Extend onDNSRecheck() to also re-resolve nameserver hostnames, not just search destinations. If the IP changes, the old connection is cleaned up and a new one is established immediately.
startNS() bailed early when nameServers was empty, leaving dnsRecheckTimer dormant. This made onDNSRecheck()'s searchDest re-resolution path dead code — hostnames from EPICS_PVA_ADDR_LIST were resolved once at startup but never re-checked. Now startNS() also checks whether any searchDest entry carries a hostname and starts the DNS recheck timer in that case. The TCP nameserver reconnect timer (nsChecker) remains gated on nameServers being non-empty.
Connection::build() inserts the new connection into connByAddr keyed by address. Assigning the result to ns.conn then dropped the last reference to the old connection, whose destructor runs cleanup() and erases connByAddr[peerAddr]. For a stable address (e.g. a Kubernetes ClusterIP that is unchanged across pod restarts), this erased the freshly inserted entry, leaving search replies unable to find the nameserver connection. Reset ns.conn before build() so the old destructor's erase runs while the dead entry is still mapped, letting the fresh insert survive. Applied to both onNSCheck() and onDNSRecheck().
…name servers self.tcp_port==0 fallback checked self.nameServers before it was populated by EPICS_PVA_NAME_SERVERS parsing, so it never fired; a bare hostname without an explicit port got resolved and cached with port 0, breaking later DNS re-resolution keying.
…change onNSCheck was reconnecting hostname-based nameservers to potentially stale IPs (OS DNS cache) every 10s, while onDNSRecheck only rebuilt the connection when the resolved IP changed. When a k8s service is deleted+recreated with a new IP, DNS may still return the old IP during propagation — so onDNSRecheck never triggered a reconnect, and onNSCheck kept hammering the dead old IP. Fix by: - onNSCheck skips hostname entries entirely; onDNSRecheck owns them - onDNSRecheck reconnects when IP changed OR conn is Disconnected, so a downed connection always retries even if DNS hasn't propagated yet - Reduce dnsRecheckInterval 30s -> 10s to match tcpNSCheckInterval and recover within typical k8s DNS propagation time
Raises the connect/reconnect log lines in startNS/onNSCheck/onDNSRecheck from debug to info, and includes the hostname on connect, so nameserver contact attempts are visible without enabling debug logging.
Keeps the added hostname text in the log messages but reverts the level back to debug, undoing the info bump from the prior commit.
Debugging why EPICS_PVA_NAME_SERVERS hostname association is lost at runtime (reconnects go through the bare-IP onNSCheck path instead of onDNSRecheck). Traces hostnameMap population in split_addr_into and the lookup in ContextImpl's nameServers build loop.
… resolution setAddress() copied the pre-colon substring into a fixed scratch[INET6_ADDRSTRLEN+1] (47-byte) buffer and threw "IPv4 address too long" whenever it overflowed, before ever reaching the evutil_inet_pton/GetAddrInfo fallback used for hostname resolution. Kubernetes Service FQDNs (e.g. foo-service.namespace.svc.cluster.local) routinely exceed 46 characters, so any such hostname given via EPICS_PVA_NAME_SERVERS was rejected outright instead of being resolved, silently defeating the DNS-recheck/hostname-tracking reconnect logic. Use a std::string instead of a fixed buffer to remove the length limit.
epicsEnvUnset() is only available on newer EPICS base (7.0+) and broke the build on the "Native Linux with 3.14" CI job. Add a small unsetEnv() helper using unsetenv()/_putenv() instead.
Contributor
Author
|
@mdavidsaver what do you think about this behaviour? |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PVXS Hostname Re-resolution and Name Server Hostname Support
Summary
This change allows PVXS clients to keep working when a configured hostname starts resolving to a different IP address. It also allows
EPICS_PVA_NAME_SERVERSto contain a hostname instead of requiring a literal IP address.The client periodically resolves hostname entries again. When the resolved IP changes, PVXS uses the new address for searches or name-server connections without requiring a client restart.
Scope
This change covers the following configuration variables:
EPICS_PVA_ADDR_LISTaccepts literal IP addresses and hostnames. Hostname entries are retained after startup and periodically resolved again for UDP searches.EPICS_PVA_NAME_SERVERSaccepts literal IP addresses and hostnames. Hostname entries are periodically resolved again for TCP connections.Literal IP addresses keep their existing behavior and are not resolved again. Hostname entries are checked approximately every 10 seconds. For
EPICS_PVA_ADDR_LIST, the initial address and the original hostname are both kept. If the hostname later resolves to a different IP, the search destination is updated and subsequent search retries use the new address. ForEPICS_PVA_NAME_SERVERS, a changed address causes the old connection to be replaced with a connection to the new address. Disconnected name-server connections are also retried during this check.Previous Behavior
EPICS_PVA_NAME_SERVERSaccepted only literal IP addresses. A hostname in this variable could not be configured successfully.Hostname entries in the supported client configuration were resolved when the client started and kept using that address. If the hostname later resolved to a different IP, PVXS continued sending searches or connection attempts to the old address.
Configuration Examples
Both variables can use a literal IP address or a hostname:
The hostname must reach PVXS unchanged. A wrapper or container entrypoint that converts it to a literal IP before starting PVXS prevents periodic re-resolution.
Reproduce the Issue
Use the test repository epics-k8s-reconnection-test to run the test in a local Kubernetes cluster.
Prerequisites
Install the following tools before starting:
The test repository checks for
kind,kubectl,docker, andhelmwhensetup-cilium.shstarts.git clone https://github.com/slaclab/epics-k8s-reconnection-test.git cd epics-k8s-reconnection-testBuild and start the Cilium cluster. To reproduce the issue with the old behavior, use the repository defaults:
it should print something like:
the script will start building
epicsandpvxsusing the PVXS_REPO and PVXS_BRANCH. At the end verify that all monitor pods are receiving updates before changing the services:kubectl -n epics-test get pods kubectl -n epics-test logs -f \ -l 'app in (monitor-ca,monitor-pva,monitor-pvxs,monitor-pvxs-ns)' \ --prefix=trueIn another terminal, delete and recreate the services with the repository helper. This assigns new IP addresses to the services:
After the services have new IP addresses, restart both PVXS IOC deployments:
With the old PVXS behavior,
monitor-pvxsandmonitor-pvxs-nsstop receiving updates because they continue using the addresses resolved before the service recreation.To verify the fix, rebuild the test repository with the patched PVXS repository and branch:
Repeat the service recreation and IOC restart steps. Both PVXS monitors should resume receiving updates after the hostname entries are checked again.
To clean test environment: