1. Database management
Production environments can generate large amounts of transient data, as shown in the Data Rate section. Transient data is automatically purged.
For information on configuring how transient data is purged, see Purge Transient Data. For details on the transient data purge command in the Command Line Interface, see Transient Data Purge.
2. Determining server and database health
Load balancers route traffic to an available server and require a mechanism to determine if the server is available or is offline. This is known as a health check. Load balancers typically run a health check by sending a request to a destination and examining the response.
Digipass S3 provides REST endpoints for its Authentication Server, Admin Server, and API Server to check if these servers, as well as the databases that they use, are accessible. These endpoints can be accessed via an HTTP GET method. Use the response from these endpoints to determine the condition of the Digipass S3 Servers and then update the load balancer routing table. See Setting up Health Checks.
Recommendation
If you have a mission-critical server or are experiencing problems in the connection between your Digipass S3 Servers and databases, you can run the Digipass S3 health check every few seconds or every few minutes.
3. Encryption keys and rotation
The Digipass S3 Server uses symmetric encryption based on an encryption key generated at the time of tenant creation to encode temporary state information needed for protocol operations. The Digipass S3 Server provides an option to generate new encryption keys which satisfies key rotation requirements. For more information see Rotate Encryption Keys.
Recommendation
If your company policy requires rotation of encryption keys, you can schedule creating a new encryption key to replace one in current use.
4. Log management
The Digipass S3 Server provides two types of logging facilities: the diagnostic log and the audit log.
4.1. The diagnostic log
The diagnostic log (nnl.log) is an invaluable tool for troubleshooting issues during deployment. When enabled, however, it can quickly grow in size, consuming system resources. Select an appropriate logging level for your needs that balances the amount of information captured with the troubleshooting you have to do. You can use a higher log setting, like INFO, during initial and/or pre deployment. During the production rollout phase, adjust the log level down to ERROR to reduce the performance overhead of using a higher log level.
See Configure Diagnostic Logs for instructions on how to configure the Digipass S3 diagnostic logs.
Recommendations
Use a log level of ERROR during production.
Select an appropriate logging level for your needs that balances the amount of information captured with the troubleshooting requirements.
4.2. The audit log
Audit logging provides a way to help meet regulatory compliance. Audit logs record operation data for runtime authentication and registration operations performed by users, as well as operations performed by administrators. Audit logs are set up to record the finish registration and finish authorization operations.
The Digipass S3 audit log is a tamper-evident log. Running a checksum on the logs with the nnl-mgmt.sh tool provides high confidence that no one has tampered with the logs after creation.
Consider moving older audit logs to archival storage for long-term retention. This ensures that disk usage by audit logs doesn't keep growing.
See section Configure Audit Logs for more details about the audit log.
Recommendation
Move older audit logs to archival storage to control disk usage by audit logs.
5. Clean up inactive registrations
During normal operation, some number of inactive registrations accumulate. An inactive registration might exist because:
a user created a registration, but then deleted the app.
a user created a registration, but didn’t use it.
a user acquired a new device and created a new registration.
For these and other cases, you can identify all inactive registrations and disable them, then physically delete them from the database.
For example, to identify, disable, and purge all registrations that haven't been used for one year:
./nnl-mgmt.sh authenticators disable -inactivity 365
./nnl-mgmt.sh authenticators purgeRecommendation
Purge inactive registrations once a quarter. To minimize impact on system performance, schedule this purge operation at a time when you expect a low usage.
6. Updating authenticator metadata
Vendors frequently release new authenticators or update existing ones. In addition, vulnerabilities may be discovered in existing authenticators, requiring that their use be limited or phased out. This information is stored in the FIDO Metadata Service (MDS) hosted by the FIDO Alliance.
To protect yourself against vulnerabilities in trusted authenticators and get the latest certification status for authenticators, you should fetch this metadata periodically. Data downloaded from FIDO MDS includes the date when the next update will be provided. Fetch the metadata statements by that date.
Digipass S3 supplies two scripts that you can use to fetch and then load the authenticator metadata in the Authentication Server. Please refer to Update Authenticator Metadata.
Recommendations
Fetch authenticator metadata periodically. You should fetch the metadata statements by the date of the next update indicated by the FIDO MDS.
Review the report generated by the metadata fetch script.
Determine new authenticators you want your end-users to use. Import the metadata for those authenticators and update your FIDO policies if necessary.
Identify authenticators you want to discontinue due to security compromises. Update your FIDO policies to remove these authenticators. In addition, disable the metadata for those authenticators.
7. Monitoring
The guidelines described in this section are general because specific recommendations depend on the deployment architecture as well as the platform you are using.
Monitor CPU and memory usage in your Digipass S3 Servers as well as disk utilization in the Operational Databases so you can prepare for increased customer usage. Avoid situations where utilization spikes too quickly because this increases latency which negatively impacts your end users. Ideally, you should have autoscaling in place to automatically increase resources up to a set limit per your company's policies.
In addition to resources such as CPU, you should monitor the health of your cloud server instances and load balancers. The load balancers can use Digipass S3's Authentication Server, Admin Server, and API Server health check endpoints to check if these servers, as well as the databases that they use, are accessible. See Setting up Health Checks.
Registrations and authentications are key indicators for Digipass S3 Software. Look for trends in the total number of registration and authentication attempts and failures by using a tool of your choice, such as Splunk, to monitor nnl.log, the Auth Server's log file. Increased failures could be linked to rolling out a new app, how an app is using FIDO, upgrading a product, or a configuration error.
It's a good idea to monitor scheduled cron jobs. The Digipass S3 transient data purge is run on a regular basis but should also be monitored for proper functioning.
You can verify that the transient data purge is working in one of the following ways:
Check nnl-admin.log, the purge thread logs its progress in this log file.
Run the following query every hour to get the total record count in the transient_data table in the Operational Database:
select count(*) from transient_data;
If the purge thread is working properly, the count goes up and down over a period of time. Otherwise, it keeps increasing.
The type of monitoring you can do depends on the available tools. Some tools allow you to set a threshold for usage and a polling frequency. For example, your company could decide to set an 80% CPU utilization threshold that's polled every minute. Once that threshold is crossed, you have several strategies at your disposal. You could automatically trigger autoscaling to respond. You could implement scripts to execute tasks. Or you may choose to use an escalation strategy to alert a human operator to the situation. The monitoring system could begin by first sending email then SMS messages with increased frequency.
Alternatively, you might have tools, such as Prometheus or Grafana, at your disposal that enable you to proactively query your servers and databases for the status of their CPU, memory, and other resources. This data can be graphed over time to look for trends. You need to balance how much data to collect, how frequently to get data, and how long to store the data versus the costs to retrieve and store that data.
Recommendations
Monitor the following:
CPU utilization
Memory usage
Disk utilization
Server health using the health check endpoints of Digipass S3 server components
Load balancer pool distribution statistics
Total number of registration and authentication attempts and failures
Digipass S3's transient data purge
cron jobs
Ideally, have autoscaling in place to automatically increase resources when load or utilization increases. Otherwise, ensure that the right personnel are notified about impending issues.
If your monitoring tools examine data over time, balance the cost to retrieve and store that data versus how much data to collect, how frequently to collect that data, and how long to store the data.
7.1. Slow response warning log
In order to detect slow responses to REST API calls, configure the elapsed time threshold system property nnl.api.elapsed.time.threshold.millis. If the response time to a REST API operation request is longer than this property's value (default: 3000 milliseconds), the Server logs a warning in the nnl.log file. Below is an example that was logged after a response to an INIT_ADAPTIVE request that took longer than the threshold that is set to 1000 milliseconds:
2024-08-14 14:50:09,729 [http-nio-8443-exec-1] [tenantId : default] [CorrelationId : jR0ee5E024td_6_5sAw7WA] com.noknok.unified.rest.v2.metrics.MetricsCollector populateElapsedTimeForResponse - WARN: ID: [jR0ee5E024td_6_5sAw7WA], Operation: [INIT_ADAPTIVE] Elapsed time: [1606 milliseconds]. Exceeds threshold: [1000 milliseconds].8. Server updates
When you need to upgrade the Digipass S3 product on a server, apply updates in the following order:
Database schema updates
Authentication Server updates
API Server updates
For zero-downtime, update Servers in a multisite configuration. This also provides the most options if you need to revert the update. When updating Servers, update one site only and stop database replication traffic until the update is tested successfully. Once the first site is successfully updated, update the second site and re-enable database replication between sites. If an update needs to be reverted, traffic can still flow through the non-updated site while the updated site is reverted.
After upgrading, end users can continue to use apps developed with the previous release. However, they cannot use new features included in the upgrade.
In the diagrams below, the NNL server icon represents the entire Digipass S3 Server cluster that includes the Authentication, API, Admin, and command-line-tools servers.
Successful update flow - Active, Standby Configuration
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 1 Initial state
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 2 Stop replication between sites
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 3 Route traffic through site B and update site A
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 4 Once site A is updated, route traffic through site A and test
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 5 Once site A testing is successful, re-enable DB replication and update site B
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 6 Re-enable traffic through site B
Unsuccessful update and recovery flow
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 7 Initial state
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 8 Stop replication between sites
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 9 Route traffic through site B and update site A
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 10 Once site A is updated, route traffic through site A and test
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 11 Site A testing is unsuccessful, send traffic back to site B
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 12 Restore site A authentication servers back to original version and restore site A database servers
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 13 Re-enable DB replication between sites, data will flow from site B to site A
.png?sv=2026-02-06&spr=https&st=2026-09-30T03%3A53%3A49Z&se=2026-09-30T04%3A12%3A49Z&sr=c&sp=r&sig=CLVjewfAc5k2V1X%2BT1aktC06t3nooM3U1WSSVZm8Swk%3D)
Figure 14 Test flow through site A. If successful, route traffic through site A
Recommendation
When updating Servers, initially update only one site and stop database replication traffic until the update is tested successfully. Once the first site is successfully updated, update the second site and re-enable database replication between sites.