Elixir for IoT: Building Platforms That Keep Working

IoT demo can look simple: a sensor sends a reading, the backend saves it, and a dashboard updates. The real test comes later. A site loses connectivity. Thousands of devices reconnect together. Old readings arrive after new ones. Operators need to know whether a device is offline or simply has nothing to report.
These are the problems that make Elixir a perfect choice for IoT. Its strengths fit systems that manage many connections, coordinate ongoing work, and recover from failures. With the right architecture, it can support everything from an edge gateway to the backend and the operator dashboard.
This article explains where Elixir helps, how the main tools fit together, and which decisions still need careful engineering.
Table of contents
- Why Elixir works well for IoT
- What a production IoT architecture looks like
- How to handle bursts of device data
- When to run Elixir on the device
- How to keep a device fleet healthy
- Where MQTT, Kafka, and RabbitMQ fit
- How to move from prototype to production
- Elixirator’s IoT experience
- Key takeaways
- Elixirator’s Elixir IoT development services
- Frequently asked questions
- Useful materials to read
Why Elixir works well for IoT
Elixir runs on the BEAM, the Erlang virtual machine. Its lightweight processes let an application handle many independent activities at once: tracking connections, receiving readings, managing timeouts, or sending commands.
A process can represent an active device session or a gateway. This is a useful design option, not a requirement to keep one process running for every registered device.
Recovery is part of the design
OTP supervisors monitor processes and restart them according to a defined strategy. For example, a failing connection handler can restart without requiring the whole application to restart.
That makes recovery easier to organize. It does not remove shared failure points: a database outage, exhausted memory, or a poorly designed supervision tree can still affect many devices.
The practical benefit is control over how the application reacts when something breaks.
Distribution gives you useful building blocks
Distributed Erlang lets processes communicate across connected nodes. That helps when work needs to run across several servers.
Your team still has to decide which node owns a device session, how state is recovered, and what happens when nodes lose contact. Communication between nodes does not automatically replicate application state or resolve conflicting updates.
A shared language can simplify development
Using Elixir across an embedded application, backend, and interface can reduce context switching and make shared logic easier to reuse.
A 2024 Elixir/Potato research paper reported 75 lines of code for its tierless implementation versus 562 for the compared Python implementation—about 87% fewer. That is a result from a specific research example, not a prediction for every IoT project or proof of fewer bugs.
For a product team, the more useful question is whether a shared stack makes its own system easier to build and maintain.
What a production IoT architecture looks like

Start by separating responsibilities. Device communication, business rules, storage, and dashboards have different needs.
| Layer | Responsibility | Possible tools |
|---|---|---|
| Device or gateway | Read sensors, buffer data, communicate with the backend | Existing firmware, Nerves, Circuits |
| Connectivity | Authenticate devices and exchange messages | MQTT broker, HTTP, WebSockets |
| Processing | Validate events and apply business rules | Elixir/OTP, Broadway or GenStage where useful |
| Storage | Keep device records, readings, and command history | PostgreSQL; additional stores when justified |
| Operator interface | Show fleet status and support actions | Phoenix LiveView |
This is a starting point, not a list of components every platform needs.
For example, a temperature-monitoring service might receive readings through MQTT, validate them in Elixir, save them, and update an operator’s view. A local gateway could buffer readings during an outage and upload them when connectivity returns.
Separate live state from durable records
Process memory and ETS can hold useful working state, such as a recent reading or an active connection. ETS is an in-memory store; it is not a replacement for durable storage. See the Elixir guide to Erlang libraries.
Decide what must survive a restart: device identity, alert history, pending commands, or billing events. Store those records durably and define how the application reconstructs its working state.
For telemetry storage, evaluate write volume, query patterns, retention, and cost. A time-series or analytical database may help with a particular workload, but there is no universal device-count threshold at which PostgreSQL stops being suitable.
Keep the dashboard useful under load
Phoenix LiveView maintains server-side UI state and sends changes to connected browsers. It can support an interactive fleet dashboard without requiring a separate frontend framework, although some interfaces still need custom JavaScript.
Design the view around what operators need to act on. A summary of connection status and active alerts is often more useful than rendering every incoming reading. Grouping updates also helps prevent the interface from becoming the next bottleneck.
How to handle bursts of device data
Average traffic can hide the hardest part of an IoT workload.
Imagine 10,000 devices reporting once a minute. If their reports are evenly spread, that is about 167 messages per second. If their timers align, the backend may receive most of those messages together. After an outage, it may also receive a backlog.
Plan for those conditions before choosing a target throughput.
Broadway provides concurrent processing, batching, acknowledgements, and demand-based backpressure through GenStage. GenStage is useful when you need more control over a custom pipeline.
Broadway does not provide retries out of the box or act as durable message storage. Retry behavior depends on the producer, the source system, and your configuration. Its backpressure also does not automatically slow every device upstream.
A practical ingestion design should answer five questions:
- Where can messages wait? Use bounded buffers and durable storage where loss is unacceptable.
- What happens to duplicates? Give events stable identifiers and make repeated processing safe.
- Does order matter? Use device sequence numbers or versions where needed; concurrent processing does not preserve order by default.
- What happens after repeated failure? Define retry limits and a place to inspect failed events.
- Can commands expire? A delayed command may no longer be appropriate when a device reconnects.
Test these behaviors with realistic payloads and outage scenarios, not only a steady stream of small messages.
When to run Elixir on the device
Elixir can run in the backend while devices use entirely different firmware. Running it at the edge is an additional option.
Use Nerves for supported embedded hardware
Nerves builds an embedded software image containing the Linux components and Elixir application your device needs. It is useful for gateways and controllers on supported Linux-capable hardware, including Raspberry Pi targets.
Check hardware support, memory, startup requirements, storage, and power consumption before choosing it for a product.
Use Circuits to access hardware
Circuits provides libraries for interfaces such as GPIO, I²C, SPI, and UART.
Nerves and Circuits work together: Nerves builds the device system, while Circuits lets the application communicate with peripherals. Circuits can also be used in other compatible environments.
Keep strict timing in the appropriate hardware layer
For a small microcontroller or a control loop with hard timing deadlines, use firmware and hardware designed for that job. Elixir can coordinate the wider workflow through a gateway or backend.
This division lets you keep low-level control close to the hardware while using Elixir for communication and supervision.
How to keep a device fleet healthy

A responsive API does not prove that devices are reporting correctly. Monitor the full journey from the device to the operator.
Useful measurements include:
- Time since the last reading, compared with the device’s expected reporting interval.
- Time from receiving an event to processing it.
- Queue depth, oldest queued message, and retry volume.
- Command acknowledgements and expired commands.
- Process mailbox growth, memory use, and scheduler load.
- Firmware versions and update failures.
The Telemetry library provides a common way to emit and handle instrumentation events. Your application and monitoring setup still need to collect the measurements that matter.
Test failure recovery deliberately: disconnect devices, slow a dependency, restart a node, and restore a site with buffered readings. Check both whether service recovers and whether the resulting data is correct.
Treat backend releases and firmware updates separately
OTP supports hot code upgrades, but they require release and state-transition planning. They are not the same as replacing a device’s firmware image.
For backend releases, a rolling deployment can be a practical starting point. Allow work to finish where possible, manage reconnections, and test recovery from interrupted processing. A grace period alone does not guarantee that no events are lost.
For Nerves fleets, NervesHub provides firmware update and device management capabilities. Plan small initial rollouts, health checks, signature verification, and a tested recovery path. Assume a device may lose power or connectivity during the update.
Where MQTT, Kafka, and RabbitMQ fit
These tools solve different problems. Choose them by responsibility rather than treating them as interchangeable upgrades.
| Tool | Main role | Key consideration |
|---|---|---|
| Elixir/OTP | Connection handling, device state, application logic | In-memory state and messages need an explicit recovery strategy |
| MQTT | Lightweight publish/subscribe communication with devices | Configure sessions, delivery quality, authentication, and broker persistence |
| Kafka | Retained event streams, replay, and independent consumers | Partitioning, retention, and consumer behavior need planning |
| RabbitMQ queues | Routing messages and distributing work | Choose queue durability, acknowledgements, and failure handling deliberately |
| RabbitMQ Streams | Retained logs that consumers can read again | Streams have different semantics from conventional queues |
MQTT is useful for constrained clients and unreliable networks. Its delivery quality settings are part of the transport design; they do not replace checking whether a business operation completed successfully.
Kafka can make sense when several systems need the same event stream or need to reprocess retained events. You do not have to wait until a database reaches its limits to have that requirement.
RabbitMQ Streams also support retention and replay. Conventional queues normally remove messages after successful acknowledgement, so the distinction matters.
Keep the architecture as small as the requirements allow. A platform may need MQTT and PostgreSQL without needing a separate event-streaming cluster.
How to move from prototype to production
A useful first milestone proves one complete workflow, including what happens when it fails.
| Stage | What to establish | Useful output |
|---|---|---|
| Discovery | Hardware constraints, traffic patterns, latency needs, and existing integrations | Architecture proposal and main risks |
| Proof of concept | A working path from device event to stored record and dashboard | Demonstration and initial load-test results |
| Pilot | Behavior across representative devices and network conditions | Monitoring, recovery procedures, and update process |
| Production rollout | Capacity, operational ownership, and controlled expansion | Measured service targets and support responsibilities |
Treat this as a planning framework. Timelines depend on hardware readiness, integrations, access to real devices, and the work already completed.
Before scoping the project, gather the expected number of connected devices, normal and peak message rates, payload sizes, acceptable delays, offline behavior, and retention needs. Include an example of a command or alert that must not be lost or applied twice.
Elixirator’s IoT experience
StorageDefender is a practical example of Elixirator’s work with connected devices.
Its platform provides wireless monitoring for self-storage facilities. The published case study describes work on high-throughput messaging, clustered services using Kubernetes and Mnesia, audit logs, role-based access control, and third-party integrations.
Those are useful examples of the engineering behind an IoT product: moving device data reliably, managing access, and connecting the platform to the rest of the business.
The case study demonstrates experience with these challenges. The architecture for another fleet should still follow that product’s hardware, traffic, and operating requirements.
Key takeaways
- Elixir fits connection-heavy, event-driven IoT platforms. Its process model and supervision tools help organize concurrent work and recovery.
- Reliability requires explicit decisions. Persistence, duplicates, ordering, and offline behavior need to be designed and tested.
- The tools have distinct roles. Nerves handles embedded systems, Circuits accesses hardware, and LiveView supports operator interfaces.
- Measure the difficult conditions. Reconnection bursts and dependency failures can reveal problems that average throughput tests miss.
Elixirator’s Elixir IoT development services
Whether you are building a new platform or improving an existing one, Elixirator can help with Elixir architecture, device orchestration, real-time data processing, and Phoenix interfaces.
Our Elixir development services include consulting and dedicated engineering support. The right starting point may be a focused architecture review, a prototype, or engineers joining your team for ongoing development.
Tell us what your devices do, where the current platform struggles, and what needs to change. We can help define the next practical step.
FAQ
Is Elixir a good choice for a large IoT fleet?
Yes, particularly when the backend manages many connections and ongoing device workflows. Capacity depends on message frequency, payloads, processing work, and storage—not just the number of registered devices. Benchmark a representative workload before setting fleet-size targets.
Do all devices need to run Elixir?
No. Devices can keep their existing firmware and communicate with an Elixir backend through supported protocols. Nerves is an option for suitable gateways or embedded devices, not a requirement for adopting Elixir on the server.
Does Elixir prevent data loss automatically?
No. Supervision helps restart failed processes; it does not automatically preserve their state or replay messages. Define durable storage, acknowledgements, retries, and duplicate handling around the operations that must survive failure.
Do we need Kafka or RabbitMQ from the start?
Only if your requirements justify them. Some platforms can start with a device protocol, an Elixir application, and a database. A broker or event log becomes useful when you need buffering, independent consumers, routing, or replay beyond that design.
Can Elixir be introduced into an existing IoT platform?
Yes. A contained component—such as a connection service, processing pipeline, or operator dashboard—can be a useful first step. Define its interfaces and success criteria, then assess its behavior under realistic traffic before expanding its role.
Useful materials to read
- OTP supervisor documentation — restart strategies and supervision design.
- Distributed Erlang — communication between runtime nodes.
- Broadway documentation — processing, acknowledgements, ordering, and failure behavior.
- Nerves: Getting started — embedded systems and supported hardware examples.
- Elixir Circuits — libraries for communicating with peripherals.
- NervesHub documentation — firmware updates and fleet management.
- Phoenix LiveView documentation — interactive interfaces built with Elixir.
- MQTT overview — device messaging and delivery quality levels.
- Apache Kafka introduction — durable event streams and processing.
- RabbitMQ Streams — retention and replay, with a comparison to queues.
- The Benefits of Tierless Elixir/Potato for Engineering IoT Software — the specific research behind the code-size comparison.
- StorageDefender case study — Elixirator’s work on a production IoT platform.

