Server-side analytics in Rails, without a banner
Most analytics start with a script in the browser. For a Rails app there is a simpler place: the server already sees every page a signed-in person opens and every change they make. Recording it there needs no script, no cookie and no consent banner, cannot be blocked, and knows exactly which user and which account did what.
This is how we wire datalove into our own Rails apps. The same pattern works with any Segment-compatible endpoint.
What to record
- Signed-in screens. An
after_actionrecords a page view for every HTMLGETthat answered 200, named after the action (customers#show), never the address or its query: addresses carry ids, searches and tokens. - Every change. A second
after_actionrecords a successfulPOST,PATCH,PUTorDELETE(a 2xx, or a redirect without an alert) ascustomers#update. Nothing depends on remembering to add a call when a feature ships. - Named actions where a name says more:
track_usage("Invoice Sent", kind: "reminder")after the success boundary, never before. It replaces the generic event for that request, so nothing counts twice. - Who, at sign-in and sign-up: an
identifywith the person’s name and role and agroupwith the account’s name. Every event carries the user id and the account (context.groupId), so accounts can be analysed as accounts. - Public pages, anonymously, as described further down.
And what never to record: pages with a token in their address (password resets, invitations, links for end customers of your customers), webhooks, the API, and background requests.
Send it the way Segment’s libraries do
Recording must never slow down or break a page, and it needs no database of its own. Segment’s Ruby library already does what is needed: calls go into an in-memory queue, and a background thread sends them in batches and retries with backoff while the endpoint is unreachable. datalove speaks the same API, so the library only needs a different host:
# Gemfile
gem "analytics-ruby", require: "segment/analytics"
# app/services/product_analytics.rb
module ProductAnalytics
def self.client
@client ||= Segment::Analytics.new(write_key: ENV.fetch("DATALOVE_WRITE_KEY"), host: "www.datalove.io",
on_error: ->(status, error) { Rails.error.report(RuntimeError.new("#{status}: #{error}"), handled: true) })
end
end
# config/initializers/product_analytics.rb: send what is still queued when a deploy stops the process
at_exit { Timeout.timeout(5) { ProductAnalytics.client.flush } rescue Timeout::Error }
Every message carries a messageId, so a batch that is sent twice is stored once. What is still in memory when a
process crashes is lost: a few seconds of events, which is the trade Segment makes too. datalove keeps whatever
arrives durably. If a single event must never be lost, send that one from a background job with retries instead.
In tests, run the library with test: true: calls land in client.test_queue instead of the network, and a test
asserts what was queued.
The concern
module ProductUsage
extend ActiveSupport::Concern
included do
after_action :record_page_view
after_action :record_action
end
private
def record_page_view
return unless request.get? && response.status == 200 && request.format.html?
return if turbo_frame_request? || prefetch? || !@authentication_required
ProductAnalytics.client.page(user_id: Current.user.id.to_s, name: "#{controller_path}##{action_name}",
context: { groupId: Current.user.account_id.to_s })
rescue StandardError => error
Rails.error.report(error, handled: true) # analytics is never a reason for a 500
end
# record_action and track_usage follow the same shape with .track
end
@authentication_required is set where the Rails authentication generator’s require_authentication resumes a
session. Controllers that allow_unauthenticated_access are not app screens.
Turbo swallows clicks, unless you tell it not to
This one cost us a day of missing data. Turbo 8 prefetches a link when the pointer rests on it, and on click it shows the prefetched page. A prefetch is not a view, so the concern skips it, and the click itself never reaches the server. Every click on a link you had hovered first was lost.
<meta name="turbo-prefetch" content="false">
in every layout of the app and the public site. Navigation stays instant enough for a server-rendered app.
The other way round, pages that refresh themselves are counted as views they are not: a dashboard that polls with
Turbo.visit(location.href) during a long job, or a page that receives broadcast_refresh. Mark those requests:
let refreshing = false
document.addEventListener("turbo:before-stream-render", (event) => {
if (event.target.getAttribute("action") === "refresh") refreshing = true
})
document.addEventListener("turbo:before-fetch-request", (event) => {
if (!refreshing) return
event.detail.fetchOptions.headers["X-Page-Refresh"] = "1"
refreshing = false
})
and skip requests with X-Page-Refresh on the server. Turbo frames (Turbo-Frame header) are parts of a page and
skipped too.
Public pages without cookies
For visitors who are not signed in, the server can still count unique visitors per day, the way Plausible does: a SHA-256 of a random salt that is replaced and deleted every day, the host, the address and the user agent. The article on counting visitors without cookies explains why that needs no banner. From the same request, derive the country, region and city (a local DB-IP Lite database) and the browser, OS and device type, and store only those, never the address or the user agent string.
Leave out known bots by user agent, and requests without one. Record the referring site’s domain, never its address,
and the utm_* parameters. A sign-up can be sent as an anonymous goal of the day’s visitor, separate from the
person’s own events, so sources can be compared by the sign-ups they bring without tying a visit to a person.
Tell people, and let them object
Describe what is recorded, why (legitimate interest, Art. 6 (1) (f) GDPR) and for how long in your privacy notice.
Keep one column, users.product_events_objected_at, that switches recording off for a person who asks, and erase
what was recorded. This is our reading of the rules for B2B software, not legal advice.
Questions and answers
How much does it slow down a request? Almost nothing. Queueing a call in memory took 0.008 ms on a laptop (5,000 calls measured); the sending happens on a background thread. For public pages add the visitor hash (0.06 ms, one indexed lookup of the day’s salt), the DB-IP lookup (0.17 ms, read from the file, not loaded into memory) and parsing the user agent (0.12 ms): about 0.4 ms. Measure it on your own hardware before you trust ours.
Why not a table in our database, to be safe? We tried that first: an outbox table and a job. It never loses an event, but it adds a migration, a write on every request and a job to every app, to protect a few seconds of analytics that only matter in a crash. The endpoint’s own storage is durable; the queue in memory only has to bridge seconds.
What if datalove is down? The library keeps calls in memory (up to 10,000 by default) and retries with backoff. A short outage costs nothing; a long one drops the oldest calls once the queue is full, and your pages never notice.
What about deploys?
A deploy stops the old process gracefully; the at_exit hook sends what is queued, for at most five seconds, so a
deploy never hangs on an unreachable endpoint.
Can a batch arrive twice?
Yes, when an answer is lost and the library retries. Each message has its own messageId, and datalove stores a
message id once.
Does the browser run anything? No, apart from the few lines that mark refreshes. Nothing is stored in the browser, which is also why no consent banner is needed for it.
What about single-page apps or pages that change without a request? Server-side recording sees requests. Interactions that never reach the server (a tab switch in the browser, a chart hover) are not recorded; if one matters, make it a request or send a named event from the action that follows.
How do we test it? Integration tests that make requests and assert what was queued (test mode): a screen, a refused change (nothing recorded), a prefetch, a refresh header, a bot, a page with a token. One test runs the real client against a WebMock stub of the endpoint and checks the batch, the write key and the message id.