Elixir for IoT: Building Platforms That Keep Working

Elixir for IoT: Building Platforms That Keep Working article image cover

IoT demo can look simple: a sensor sends a reading, the backend saves it, and a dashboard updates. The real test comes later. A site loses connectivity. Thousands of devices reconnect together. Old readings arrive after new ones. Operators need to know whether a device is offline or simply has nothing to report.

These are the problems that make Elixir a perfect choice for IoT. Its strengths fit systems that manage many connections, coordinate ongoing work, and recover from failures. With the right architecture, it can support everything from an edge gateway to the backend and the operator dashboard.

This article explains where Elixir helps, how the main tools fit together, and which decisions still need careful engineering.

Table of contents

Why Elixir works well for IoT

Elixir runs on the BEAM, the Erlang virtual machine. Its lightweight processes let an application handle many independent activities at once: tracking connections, receiving readings, managing timeouts, or sending commands.

A process can represent an active device session or a gateway. This is a useful design option, not a requirement to keep one process running for every registered device.

Recovery is part of the design

OTP supervisors monitor processes and restart them according to a defined strategy. For example, a failing connection handler can restart without requiring the whole application to restart.

That makes recovery easier to organize. It does not remove shared failure points: a database outage, exhausted memory, or a poorly designed supervision tree can still affect many devices.

The practical benefit is control over how the application reacts when something breaks.

Distribution gives you useful building blocks

Distributed Erlang lets processes communicate across connected nodes. That helps when work needs to run across several servers.

Your team still has to decide which node owns a device session, how state is recovered, and what happens when nodes lose contact. Communication between nodes does not automatically replicate application state or resolve conflicting updates.

A shared language can simplify development

Using Elixir across an embedded application, backend, and interface can reduce context switching and make shared logic easier to reuse.

A 2024 Elixir/Potato research paper reported 75 lines of code for its tierless implementation versus 562 for the compared Python implementation—about 87% fewer. That is a result from a specific research example, not a prediction for every IoT project or proof of fewer bugs.

For a product team, the more useful question is whether a shared stack makes its own system easier to build and maintain.

What a production IoT architecture looks like

Elixir connecting an IoT sensor to processing, storage, and an operator dashboard

Start by separating responsibilities. Device communication, business rules, storage, and dashboards have different needs.

LayerResponsibilityPossible tools
Device or gatewayRead sensors, buffer data, communicate with the backendExisting firmware, Nerves, Circuits
ConnectivityAuthenticate devices and exchange messagesMQTT broker, HTTP, WebSockets
ProcessingValidate events and apply business rulesElixir/OTP, Broadway or GenStage where useful
StorageKeep device records, readings, and command historyPostgreSQL; additional stores when justified
Operator interfaceShow fleet status and support actionsPhoenix LiveView

This is a starting point, not a list of components every platform needs.

For example, a temperature-monitoring service might receive readings through MQTT, validate them in Elixir, save them, and update an operator’s view. A local gateway could buffer readings during an outage and upload them when connectivity returns.

Separate live state from durable records

Process memory and ETS can hold useful working state, such as a recent reading or an active connection. ETS is an in-memory store; it is not a replacement for durable storage. See the Elixir guide to Erlang libraries.

Decide what must survive a restart: device identity, alert history, pending commands, or billing events. Store those records durably and define how the application reconstructs its working state.

For telemetry storage, evaluate write volume, query patterns, retention, and cost. A time-series or analytical database may help with a particular workload, but there is no universal device-count threshold at which PostgreSQL stops being suitable.

Keep the dashboard useful under load

Phoenix LiveView maintains server-side UI state and sends changes to connected browsers. It can support an interactive fleet dashboard without requiring a separate frontend framework, although some interfaces still need custom JavaScript.

Design the view around what operators need to act on. A summary of connection status and active alerts is often more useful than rendering every incoming reading. Grouping updates also helps prevent the interface from becoming the next bottleneck.

How to handle bursts of device data

Average traffic can hide the hardest part of an IoT workload.

Imagine 10,000 devices reporting once a minute. If their reports are evenly spread, that is about 167 messages per second. If their timers align, the backend may receive most of those messages together. After an outage, it may also receive a backlog.

Plan for those conditions before choosing a target throughput.

Broadway provides concurrent processing, batching, acknowledgements, and demand-based backpressure through GenStage. GenStage is useful when you need more control over a custom pipeline.

Broadway does not provide retries out of the box or act as durable message storage. Retry behavior depends on the producer, the source system, and your configuration. Its backpressure also does not automatically slow every device upstream.

A practical ingestion design should answer five questions:

  1. Where can messages wait? Use bounded buffers and durable storage where loss is unacceptable.
  2. What happens to duplicates? Give events stable identifiers and make repeated processing safe.
  3. Does order matter? Use device sequence numbers or versions where needed; concurrent processing does not preserve order by default.
  4. What happens after repeated failure? Define retry limits and a place to inspect failed events.
  5. Can commands expire? A delayed command may no longer be appropriate when a device reconnects.

Test these behaviors with realistic payloads and outage scenarios, not only a steady stream of small messages.

When to run Elixir on the device

Elixir can run in the backend while devices use entirely different firmware. Running it at the edge is an additional option.

Use Nerves for supported embedded hardware

Nerves builds an embedded software image containing the Linux components and Elixir application your device needs. It is useful for gateways and controllers on supported Linux-capable hardware, including Raspberry Pi targets.

Check hardware support, memory, startup requirements, storage, and power consumption before choosing it for a product.

Use Circuits to access hardware

Circuits provides libraries for interfaces such as GPIO, I²C, SPI, and UART.

Nerves and Circuits work together: Nerves builds the device system, while Circuits lets the application communicate with peripherals. Circuits can also be used in other compatible environments.

Keep strict timing in the appropriate hardware layer

For a small microcontroller or a control loop with hard timing deadlines, use firmware and hardware designed for that job. Elixir can coordinate the wider workflow through a gateway or backend.

This division lets you keep low-level control close to the hardware while using Elixir for communication and supervision.

How to keep a device fleet healthy

Elixir organizing bursts of device messages into a controlled processing pipeline

A responsive API does not prove that devices are reporting correctly. Monitor the full journey from the device to the operator.

Useful measurements include:

  • Time since the last reading, compared with the device’s expected reporting interval.
  • Time from receiving an event to processing it.
  • Queue depth, oldest queued message, and retry volume.
  • Command acknowledgements and expired commands.
  • Process mailbox growth, memory use, and scheduler load.
  • Firmware versions and update failures.

The Telemetry library provides a common way to emit and handle instrumentation events. Your application and monitoring setup still need to collect the measurements that matter.

Test failure recovery deliberately: disconnect devices, slow a dependency, restart a node, and restore a site with buffered readings. Check both whether service recovers and whether the resulting data is correct.

Treat backend releases and firmware updates separately

OTP supports hot code upgrades, but they require release and state-transition planning. They are not the same as replacing a device’s firmware image.

For backend releases, a rolling deployment can be a practical starting point. Allow work to finish where possible, manage reconnections, and test recovery from interrupted processing. A grace period alone does not guarantee that no events are lost.

For Nerves fleets, NervesHub provides firmware update and device management capabilities. Plan small initial rollouts, health checks, signature verification, and a tested recovery path. Assume a device may lose power or connectivity during the update.

Where MQTT, Kafka, and RabbitMQ fit

These tools solve different problems. Choose them by responsibility rather than treating them as interchangeable upgrades.

ToolMain roleKey consideration
Elixir/OTPConnection handling, device state, application logicIn-memory state and messages need an explicit recovery strategy
MQTTLightweight publish/subscribe communication with devicesConfigure sessions, delivery quality, authentication, and broker persistence
KafkaRetained event streams, replay, and independent consumersPartitioning, retention, and consumer behavior need planning
RabbitMQ queuesRouting messages and distributing workChoose queue durability, acknowledgements, and failure handling deliberately
RabbitMQ StreamsRetained logs that consumers can read againStreams have different semantics from conventional queues

MQTT is useful for constrained clients and unreliable networks. Its delivery quality settings are part of the transport design; they do not replace checking whether a business operation completed successfully.

Kafka can make sense when several systems need the same event stream or need to reprocess retained events. You do not have to wait until a database reaches its limits to have that requirement.

RabbitMQ Streams also support retention and replay. Conventional queues normally remove messages after successful acknowledgement, so the distinction matters.

Keep the architecture as small as the requirements allow. A platform may need MQTT and PostgreSQL without needing a separate event-streaming cluster.

How to move from prototype to production

A useful first milestone proves one complete workflow, including what happens when it fails.

StageWhat to establishUseful output
DiscoveryHardware constraints, traffic patterns, latency needs, and existing integrationsArchitecture proposal and main risks
Proof of conceptA working path from device event to stored record and dashboardDemonstration and initial load-test results
PilotBehavior across representative devices and network conditionsMonitoring, recovery procedures, and update process
Production rolloutCapacity, operational ownership, and controlled expansionMeasured service targets and support responsibilities

Treat this as a planning framework. Timelines depend on hardware readiness, integrations, access to real devices, and the work already completed.

Before scoping the project, gather the expected number of connected devices, normal and peak message rates, payload sizes, acceptable delays, offline behavior, and retention needs. Include an example of a command or alert that must not be lost or applied twice.

Elixirator’s IoT experience

StorageDefender is a practical example of Elixirator’s work with connected devices.

Its platform provides wireless monitoring for self-storage facilities. The published case study describes work on high-throughput messaging, clustered services using Kubernetes and Mnesia, audit logs, role-based access control, and third-party integrations.

Those are useful examples of the engineering behind an IoT product: moving device data reliably, managing access, and connecting the platform to the rest of the business.

The case study demonstrates experience with these challenges. The architecture for another fleet should still follow that product’s hardware, traffic, and operating requirements.

Key takeaways

  • Elixir fits connection-heavy, event-driven IoT platforms. Its process model and supervision tools help organize concurrent work and recovery.
  • Reliability requires explicit decisions. Persistence, duplicates, ordering, and offline behavior need to be designed and tested.
  • The tools have distinct roles. Nerves handles embedded systems, Circuits accesses hardware, and LiveView supports operator interfaces.
  • Measure the difficult conditions. Reconnection bursts and dependency failures can reveal problems that average throughput tests miss.

Elixirator’s Elixir IoT development services

Whether you are building a new platform or improving an existing one, Elixirator can help with Elixir architecture, device orchestration, real-time data processing, and Phoenix interfaces.

Our Elixir development services include consulting and dedicated engineering support. The right starting point may be a focused architecture review, a prototype, or engineers joining your team for ongoing development.

Tell us what your devices do, where the current platform struggles, and what needs to change. We can help define the next practical step.

FAQ

Is Elixir a good choice for a large IoT fleet?

Yes, particularly when the backend manages many connections and ongoing device workflows. Capacity depends on message frequency, payloads, processing work, and storage—not just the number of registered devices. Benchmark a representative workload before setting fleet-size targets.

Do all devices need to run Elixir?

No. Devices can keep their existing firmware and communicate with an Elixir backend through supported protocols. Nerves is an option for suitable gateways or embedded devices, not a requirement for adopting Elixir on the server.

Does Elixir prevent data loss automatically?

No. Supervision helps restart failed processes; it does not automatically preserve their state or replay messages. Define durable storage, acknowledgements, retries, and duplicate handling around the operations that must survive failure.

Do we need Kafka or RabbitMQ from the start?

Only if your requirements justify them. Some platforms can start with a device protocol, an Elixir application, and a database. A broker or event log becomes useful when you need buffering, independent consumers, routing, or replay beyond that design.

Can Elixir be introduced into an existing IoT platform?

Yes. A contained component—such as a connection service, processing pipeline, or operator dashboard—can be a useful first step. Define its interfaces and success criteria, then assess its behavior under realistic traffic before expanding its role.

Useful materials to read

elixir
iot
nerves
phoenix
real-time systems

Ready to Build with Elixirator?

Prefer a quick call?
Photo of Alex Danyliak, Client Partner at Elixirator

Alex Danyliak

Client Partner at Elixirator

“Whether it’s a one-off consulting gig or a full dedicated team, let’s chat about how Elixirator can help you build something reliable, performant, and future-proof.”

Ready to Build with Elixirator?

Prefer a quick call?
Photo of Alex Danyliak, Client Partner at Elixirator

Alex Danyliak

Client Partner at Elixirator

“Whether it’s a one-off consulting gig or a full dedicated team, let’s chat about how Elixirator can help you build something reliable, performant, and future-proof.”