# Server-side analytics in Rails, without a banner

Most analytics start with a script in the browser. For a Rails app there is a simpler place: the server already sees
every page a signed-in person opens and every change they make. Recording it there needs no script, no cookie and no
consent banner, cannot be blocked, and knows exactly which user and which account did what.

This is how we wire datalove into our own Rails apps. The same pattern works with any Segment-compatible endpoint.

## What to record

- **Signed-in screens.** An `after_action` records a page view for every HTML `GET` that answered 200, named after the
  action (`customers#show`), never the address or its query: addresses carry ids, searches and tokens.
- **Every change.** A second `after_action` records a successful `POST`, `PATCH`, `PUT` or `DELETE` (a 2xx, or a
  redirect without an alert) as `customers#update`. Nothing depends on remembering to add a call when a feature ships.
- **Named actions** where a name says more: `track_usage("Invoice Sent", kind: "reminder")` after the success
  boundary, never before. It replaces the generic event for that request, so nothing counts twice.
- **Who**, at sign-in and sign-up: an `identify` with the person's name and role and a `group` with the account's
  name. Every event carries the user id and the account (`context.groupId`), so accounts can be analysed as accounts.
- **Public pages**, anonymously, as described further down.

And what never to record: pages with a token in their address (password resets, invitations, links for end customers
of your customers), webhooks, the API, and background requests.

## Send it the way Segment's libraries do

Recording must never slow down or break a page, and it needs no database of its own. Segment's Ruby library already
does what is needed: calls go into an in-memory queue, and a background thread sends them in batches and retries with
backoff while the endpoint is unreachable. datalove speaks the same API, so the library only needs a different host:

```ruby
# Gemfile
gem "analytics-ruby", require: "segment/analytics"

# app/services/product_analytics.rb
module ProductAnalytics
  def self.client
    @client ||= Segment::Analytics.new(write_key: ENV.fetch("DATALOVE_WRITE_KEY"), host: "www.datalove.io",
                                       on_error: ->(status, error) { Rails.error.report(RuntimeError.new("#{status}: #{error}"), handled: true) })
  end
end

# config/initializers/product_analytics.rb: send what is still queued when a deploy stops the process
at_exit { Timeout.timeout(5) { ProductAnalytics.client.flush } rescue Timeout::Error }
```

Every message carries a `messageId`, so a batch that is sent twice is stored once. What is still in memory when a
process crashes is lost: a few seconds of events, which is the trade Segment makes too. datalove keeps whatever
arrives durably. If a single event must never be lost, send that one from a background job with retries instead.

In tests, run the library with `test: true`: calls land in `client.test_queue` instead of the network, and a test
asserts what was queued.

## The concern

```ruby
module ProductUsage
  extend ActiveSupport::Concern
  included do
    after_action :record_page_view
    after_action :record_action
  end

  private
    def record_page_view
      return unless request.get? && response.status == 200 && request.format.html?
      return if turbo_frame_request? || prefetch? || !@authentication_required

      ProductAnalytics.client.page(user_id: Current.user.id.to_s, name: "#{controller_path}##{action_name}",
                                   context: { groupId: Current.user.account_id.to_s })
    rescue StandardError => error
      Rails.error.report(error, handled: true) # analytics is never a reason for a 500
    end
    # record_action and track_usage follow the same shape with .track
end
```

`@authentication_required` is set where the Rails authentication generator's `require_authentication` resumes a
session. Controllers that `allow_unauthenticated_access` are not app screens.

## Turbo swallows clicks, unless you tell it not to

This one cost us a day of missing data. Turbo 8 prefetches a link when the pointer rests on it, and on click it
shows the prefetched page. A prefetch is not a view, so the concern skips it, and the click itself never reaches the
server. Every click on a link you had hovered first was lost.

```erb
<meta name="turbo-prefetch" content="false">
```

in every layout of the app and the public site. Navigation stays instant enough for a server-rendered app.

The other way round, pages that refresh themselves are counted as views they are not: a dashboard that polls with
`Turbo.visit(location.href)` during a long job, or a page that receives `broadcast_refresh`. Mark those requests:

```js
let refreshing = false
document.addEventListener("turbo:before-stream-render", (event) => {
  if (event.target.getAttribute("action") === "refresh") refreshing = true
})
document.addEventListener("turbo:before-fetch-request", (event) => {
  if (!refreshing) return
  event.detail.fetchOptions.headers["X-Page-Refresh"] = "1"
  refreshing = false
})
```

and skip requests with `X-Page-Refresh` on the server. Turbo frames (`Turbo-Frame` header) are parts of a page and
skipped too.

## Public pages without cookies

For visitors who are not signed in, the server can still count unique visitors per day, the way Plausible does:
a SHA-256 of a random salt that is replaced and deleted every day, the host, the address and the user agent. The
[article on counting visitors without cookies](/articles/count-unique-visitors-without-cookies) explains why that
needs no banner. From the same request, derive the country, region and city (a local DB-IP Lite database) and the
browser, OS and device type, and store only those, never the address or the user agent string.

Leave out known bots by user agent, and requests without one. Record the referring site's domain, never its address,
and the `utm_*` parameters. A sign-up can be sent as an anonymous goal of the day's visitor, separate from the
person's own events, so sources can be compared by the sign-ups they bring without tying a visit to a person.

## Tell people, and let them object

Describe what is recorded, why (legitimate interest, Art. 6 (1) (f) GDPR) and for how long in your privacy notice.
Keep one column, `users.product_events_objected_at`, that switches recording off for a person who asks, and erase
what was recorded. This is our reading of the rules for B2B software, not legal advice.

## Questions and answers

**How much does it slow down a request?**
Almost nothing. Queueing a call in memory took 0.008 ms on a laptop (5,000 calls measured); the sending happens on a
background thread. For public pages add the visitor hash (0.06 ms, one indexed lookup of the day's salt), the DB-IP
lookup (0.17 ms, read from the file, not loaded into memory) and parsing the user agent (0.12 ms): about 0.4 ms. Measure
it on your own hardware before you trust ours.

**Why not a table in our database, to be safe?**
We tried that first: an outbox table and a job. It never loses an event, but it adds a migration, a write on every
request and a job to every app, to protect a few seconds of analytics that only matter in a crash. The endpoint's
own storage is durable; the queue in memory only has to bridge seconds.

**What if datalove is down?**
The library keeps calls in memory (up to 10,000 by default) and retries with backoff. A short outage costs nothing; a
long one drops the oldest calls once the queue is full, and your pages never notice.

**What about deploys?**
A deploy stops the old process gracefully; the `at_exit` hook sends what is queued, for at most five seconds, so a
deploy never hangs on an unreachable endpoint.

**Can a batch arrive twice?**
Yes, when an answer is lost and the library retries. Each message has its own `messageId`, and datalove stores a
message id once.

**Does the browser run anything?**
No, apart from the few lines that mark refreshes. Nothing is stored in the browser, which is also why no consent
banner is needed for it.

**What about single-page apps or pages that change without a request?**
Server-side recording sees requests. Interactions that never reach the server (a tab switch in the browser, a
chart hover) are not recorded; if one matters, make it a request or send a named event from the action that follows.

**How do we test it?**
Integration tests that make requests and assert what was queued (test mode): a screen, a refused change (nothing
recorded), a prefetch, a refresh header, a bot, a page with a token. One test runs the real client against a WebMock
stub of the endpoint and checks the batch, the write key and the message id.
