Still Fighting Cloud Performance Issues? Learn How eToro Delivers Reliable Performance Across 2 Million Daily Trading Transactions
Transcript
Fireside Chat: How eToro Powers a 24/7 Trading Platform with Silk
Speakers: Ori Weizman — Director of Solution Architecture, EMEA, Silk Yitzchak Wahnon — DBA Team Manager, eToro
Ori Weizman: All right, I guess we’re getting started now. We can start with some quick introductions — I’ll go first. My name’s Ori Weizman, I’m the Director of Solution Architecture for EMEA. I’ve been at Silk for almost five years now — it’ll be five years in November. It’s been a lot of fun working with customers like eToro. My background is in consulting, and I’ve been in technology for more than 20 years now, which I guess makes me feel kind of old. And I’m here talking with Yitzchak. Yitzchak, why don’t you introduce yourself?
Yitzchak Wahnon: I’m the DBA team manager at eToro. I’ve been in technology even longer than Ori — about 26 years. eToro is a trading platform with millions of customers — anyone can Google us and find all the details. We’re a 24/7 platform, mainly focused on when the markets are open, serving millions of customers who trade on a social platform.
Ori Weizman: Thank you for taking the time to be here — and this is kind of my schtick, so I have to ask a question before we get started: Why did the DBA get kicked out of the bar?
Yitzchak Wahnon: I don’t know. Why?
Ori Weizman: Because he kept joining tables without permission.
Okay — dad joke delivered, I love dad jokes. Thanks again for being here, and thanks to everyone joining the webinar. Hopefully you’ll all have questions afterward so we can have a bit of a live discussion at the end. But I wanted to start, Yitzchak, by understanding more about your journey to the cloud. You mentioned eToro is a trading platform serving millions of customers, and I know you started out on-prem — eToro’s been around for about 10 years, right?
Yitzchak Wahnon: Way more than that.
Ori Weizman: Way more than that — okay. So eToro’s been around a while, and you’ve had a trading platform running on-prem. What did it look like before we came into the picture? What made you decide to move away from on-prem, what were you actually running on-prem, and how did you approach the transition to the cloud?
Yitzchak Wahnon: So, as we said, we’re a trading platform. For those who don’t know, most trading activity is focused around the New York Stock Exchange. During the night, and especially over weekends, customers place orders that can’t be processed until the market opens. So we have millions of customers placing tens of thousands — sometimes hundreds of thousands — of orders that need to be processed as quickly as possible after market open. Our infrastructure is built on SQL Server, and SQL Server’s throughput is limited by the number of CPUs available, how fast each CPU can work, and how well that’s supported by memory and disk.
We stayed on-premise mainly because, well, the company’s been around for nearly 20 years, and the cloud wasn’t mature enough for a long time. We experimented with Azure VMs at least 12 years ago, and things were very basic — I remember recreating disks on VMs multiple times just to get it right, and these weren’t even production machines. The cloud simply wasn’t up to what we needed back then, so we ran on-premise with very large machines and all-flash storage, giving us fast, sub-millisecond response times across two large physical servers.
But no matter what SLA you have on your hardware, if something goes wrong, you’re stuck waiting — the engineer shows up, a part is missing, and you wait two days to get everything back to 100%. The flexibility the cloud offers just can’t be matched by an on-premise solution. That was one of the main reasons the company made the strategic decision to move to the cloud — it offers nearly limitless scaling, especially for services and things like AKS, though less so in the database world.
Our database is large — well, large for us — almost like a monolith, since all our trading happens together, so we need a fast machine with lots of CPUs and fast storage behind it. That limits our options in the cloud. We’re based in Azure, and Azure had an offering called Elastic SAN (eSAN). We tried it, and very quickly — during the POC stage — it was clear it wasn’t going to reach the IOPS we needed. At the time we were requesting around 140,000 IOPS, and nothing in the cloud offered that kind of IOPS at sub-millisecond response time. So very quickly, the Azure-native option became a showstopper for moving the database to the cloud. And without moving the database, moving just the application layer wasn’t going to offer much benefit, given latency between cloud and on-prem of up to 20 milliseconds.
At eToro, milliseconds convert to money. If we can’t process positions fast enough, we can lose money — it puts us at enormous risk, especially when markets are highly volatile. We also deal a lot in cryptocurrency, and when we made our migration to the cloud, crypto was going wild — about a year and a half to two years ago.
Ori Weizman: So the cloud doesn’t come without risk. You were operating a certain way on-prem and looking to take advantage of better technology and more availability in the cloud — it’s much easier to test something in the cloud than to rack and stack hardware. So you wanted out of that fixed environment, but as you said, there’s real risk in moving to the cloud. If I’m hearing you right, most of that risk centered on performance — your storage solution and the VMs available to you. Were there other major concerns as you went through your analysis before coming to us?
Yitzchak Wahnon: I’d say performance isn’t entirely a risk — it’s more of a limitation. Everyone knows that in the cloud, things fail: a VM disappears, a host fails and moves, or other issues come up. We knew we’d have to account for that and build a solution resilient to those failures. Across all our systems, we’ve always aimed to avoid a single point of failure, so we build in redundancy and plan ahead — whether for a single resource failure, an availability zone failure, or even a regional failure — so we can keep running, maybe not at full capacity, but still stay up and alive.
Ori Weizman: Okay, and that’s where the POC we ran together comes in. I’ve known you for a few years now, and we’ve been running in your production environment for about a year and a half. I want to talk about the POC and how we helped you overcome this. You pointed to one of the major risks you came to us with — the cloud builds in redundancy at different layers, but you still need to take advantage of what’s available and architect your solution properly. That’s where we started to see real benefits in our collaboration: our system handles these infrastructure-level failures automatically. We tested that in the POC, and it’s been proven out in production too — when something doesn’t go as expected, recovery is seamless and doesn’t affect the system as a whole. Is that fair to say? I don’t want to put words in your mouth.
Yitzchak Wahnon: For sure. We had very close cooperation — we were very careful, and we really made Silk work hard in the POC before we were willing to put our main production systems on it. It took more than half a year before we moved production over, until we’d really tried and tested things through the kinds of Azure events that could have caused us issues. The system managed to handle all of them.
In the planning stages of the POC, one of the best parts was that Silk really sat down with us — they brought Azure technologies to our attention that we weren’t even aware of, and opened up opportunities we wouldn’t have been able to use on our own. One example: the VM series we used. We were running M-series machines, which are very expensive. Silk made us aware that an E-series was coming out — we had to wait for it, constantly checking with Microsoft on when it would be generally available, since we weren’t willing to put production on anything still in public preview.
With Silk and the E-series, we were able to save costs by moving TempDB — which we’d always kept on Azure’s ephemeral disk — onto Silk storage instead, saving the cost of the D-series disk that comes with that VM family, which adds a significant amount to each VM’s cost. So we started with the disk layer, confirmed the storage worked, then tested server by server, and it proved itself. That came with close cooperation from the Silk team, who kept bringing things to our attention that we weren’t even aware of.
We also planned the architecture together to make it resilient to failure. It’s not plug-and-play — you can’t just connect it to power and expect the storage to be there. There’s tuning to do along the way, and I’m sure our configuration isn’t exactly the same as the next customer’s. So we went live on one of our lower-tier systems first, gathered performance metrics, and Silk came back with configuration changes to improve performance, which we implemented until we reached something stable. Performance has been — I’ll admit, probably better than expected, though maybe I went in a bit pessimistic — consistently sub-millisecond, and uptime has been excellent.
When there’s an Azure issue, the Silk system automatically emails us — letting us know a C-node has been replaced, or something’s happened with a D-node — and automatically opens a support case. Once the event resolves automatically, without any human intervening, the case closes and we get a summary of what happened, confirming everything’s back in full working order.
Ori Weizman: I really want to highlight this. I’ll say up front — you made us work hard in the POC, I can attest to that directly, but it was a lot of fun going through that journey with you. I think this is indicative of a lot of customers today: your business runs on this platform, and it doesn’t run if the platform isn’t running. You want to take advantage of everything the cloud offers, but there’s volatility that needs to be addressed, and having this kind of automatic healing and support behind it matters. But more than anything, I think it’s the collaboration between our teams. Like you said, this isn’t plug-and-play — we have different configurations for different needs because it’s an agile platform, and we need to work closely with your team to land on the right solution. That’s why we approached it in stages — your demo system first, then a POC, then production.
I also want to come back to something you brought up — us pointing you toward different resources you could use, like the M-series versus the E-series. The M-series is a VM family Microsoft built to address performance for customers who need that sub-millisecond latency, but it comes at significant cost. I won’t share specific numbers, but there were meaningful cost savings — and later I’ll touch on the Total Economic Impact report Forrester put together, which interviewed eToro along with a number of our other customers and puts those results, including the cost reduction, front and center.
This is really the crux of it: we’re not just bringing a platform to help you operate in the cloud, we’re bringing cloud expertise and a partnership mentality. We want to understand your business and help move it forward, and I think that’s really what’s made us successful as an organization. To call out a couple of specifics: the E-series VMs let you save significant cost while improving performance over the M-series. And something else we did well with Microsoft — you mentioned C-nodes and D-nodes — our system has a performance layer made up of C-nodes, which are VMs in Azure, and a capacity layer made up of D-nodes, also VMs in Azure. That’s a bit of an oversimplification — I’d encourage everyone to check out our architecture white paper on our website to learn more.
As we looked at your specific implementation against our platform, we’re always trying to improve, and Microsoft is also our partner — so we collaborated with Microsoft Engineering to develop a new resource for our D-nodes called an LAOSv4, a new VM shape that cuts significant cost out of implementations. We’ve actually started a process with you to introduce this so you can capture more savings — not to get more out of the relationship, but to keep you as a happy customer and pass those savings along. We’re always looking to be more efficient and to collaborate with customers like you, and this is a great example of that — and of us actually shaping Microsoft’s ecosystem, since they developed a VM largely based on specifications from our engineering team.
That’s a great part of the relationship we’ve built. I did want to go back to eSAN, which you mentioned earlier — you tested it and it fell short pretty quickly. How long were you evaluating cloud solutions before we came into the picture? Had you just started the journey, was it a long process, had you run into problems before us?
Yitzchak Wahnon: We had tried previously and researched other suppliers — it wasn’t years of research, the technology moves quickly. There were a good few months — this was before AI — of research and conversations with different suppliers. We crossed some off the list because they were incompatible with what we needed, and then, naturally, my inclination is always to try the Microsoft/Azure-native solution first, since that means fewer parties involved when there’s a problem. But that didn’t work — it fell short quickly, and they admitted pretty fast that it just wasn’t ready. I have no idea where it stands now.
Then you came along, and once I joined the POC, you lived up to what was promised. And as you said, the close working relationship didn’t stop once the contract was signed — we’re constantly talking and brainstorming together about what’s best for our architecture. It’s an ongoing, close, cooperative relationship that’s continued to bear fruit.
Ori Weizman: Yeah, and to your point — we’ve had multiple workshops where we’ve brought you in to speak directly with our engineering team, and you know that team well by now. I was actually at your office earlier today, collaborating on some exciting things. I’ll make a short pitch for Echo, since I know it’s something you’re interested in — we’re extending our snapshot capability. Right now you can take an instant snapshot and recreate a full data set; we’re looking to go a step further and actually replicate the database. It’s not implemented at eToro yet, but we’re working together to bring it to you and get feedback from customers like you along the way.
This really has to be a partnership, otherwise it doesn’t work — we’re not selling commodity software, we’re selling something that looks and feels different for every customer, with their own specific needs. That requires tight collaboration between our organizations, and our customer success team talks with you a lot too, including around upgrades. Speaking of which — since going live with you about a year and a half ago, we’ve done a number of upgrades. I know that was an area of concern in the POC, wanting to make sure upgrades would be smooth and seamless. Can you talk about what your initial concerns were, how we addressed them, and what your experience has actually been with the upgrades we’ve done together?
Yitzchak Wahnon: Sure. SQL Server is very sensitive to any sort of I/O interruption, and as I mentioned, I/O interruptions cost eToro money. So I couldn’t accept a solution that involved any downtime for patching and upgrades — which will always eventually be needed, whether quarterly, monthly, or yearly. So we were very apprehensive going in.
With that in mind, we currently run SQL Server 2022 with Always On availability groups. I don’t upgrade our primary while it’s still the primary — we always fail over first and proceed carefully — but our secondaries are constantly in use too, scaling out reads.
When we first started, we’d be hands-on with you while upgrading an SDP —
Ori Weizman: And an SDP, for the audience, is a Silk Data Pod — an instance of our platform deployed. Sorry to cut you off there.
Yitzchak Wahnon: You see, I already speak Silk.
Ori Weizman: That’s right.
Yitzchak Wahnon: Well, any maintenance on the storage platform would be hands-on at the time — monitoring, making sure everything was okay. Now we’re in the middle of a batch of upgrades across all our resources, and we don’t even connect — it’s just done. Silk is given access by our IT team, and they run the whole upgrade. We basically just do a quick check afterward to confirm everything’s okay. So the patching and the uptime are exactly as promised — the performance is what you’d expect from a solution built to support a 24/7 system.
Ori Weizman: I really want to underline this. This is a trading platform — forget about the platform being down, if it’s even slow or disrupted, that translates directly into lost money. To say there’s a trusted relationship and a partnership here is almost an understatement — to the point where we’re running non-disruptive upgrades as an extension of your team, making sure everything is handled in a way that doesn’t put the system at risk. And that’s not easy. It’s not easy to sell software, integrate teams, and build a genuine partnership. A lot of companies will sell a product and then not stay engaged until it’s time for renewal, and I think what we do really well as an organization is treat our customers like partners, because that’s exactly what they are — and I think this conversation has shown that.
I do want to highlight a bit from the Total Economic Impact report we commissioned. Forrester — an analyst group most people are probably familiar with — sent an analyst to interview a number of our customers, as well as people within our own company (they interviewed me too), and then independently published their findings. I want to share those findings on screen.
Okay, so these are the headline numbers from what Forrester’s report uncovered. I’ll preface this by saying it’s aggregated across all the customers interviewed, so it’s not eToro-specific — eToro shares in these benefits, but we’re not publishing specific figures tied to any individual company. But when you listen to this conversation between Yitzchak and me, and think about eToro as a business and how we’ve helped it move forward, and then look at the validation from a third-party independent analyst firm — this is why we exist, this is why we succeed, and I think it’s why companies are willing to entrust their critical platforms — the systems that actually drive the business and put money in the company’s pocket — to us.
So, I think this was a great conversation. Yitzchak, I really appreciate your time — I know you’re busy, and you’ve probably seen and heard enough of me for one day since we talked a lot earlier too, so I won’t keep you any longer than I have to. Any last thoughts before we wrap up and open it to questions?
Yitzchak Wahnon: I’ll add something I just thought of. I’m a DBA manager, so strictly speaking, I don’t need to be the one sitting with Silk, the storage supplier, planning everything — our IT team is involved too. I’ll say our IT team was initially very skeptical about working with a storage supplier; they weren’t overly enthusiastic. But now, a year and a half — probably closer to two and a half years — down the line, they also find it a pleasure to work with the team. I don’t handle the upgrades myself — IT works directly with the Silk engineers — and they know they’re not walking into a task that’s going to be a headache. They know the team will connect on time and that there’s a high chance everything will run smoothly, touch wood, and that it’s genuinely a pleasure to work with. So even the IT team — and everyone knows IT people can be a bit difficult to work with, unlike DBAs, who are easy — they also enjoy the cooperation and professionalism we get.
Ori Weizman: Thank you for saying that — it’s a pleasure working with you all too, and I look forward to many more years together. Let’s open it up to questions — looks like there’s already one in the chat: “How do you deal with Azure host maintenance shutting down Silk C-nodes and D-nodes, and the data quiesce that occurs from that? Are you subscribed to the Event Grid for early warning?”
The answer is yes, we’re subscribed for early warning. With Azure, a few different things can happen: planned or scheduled maintenance, unscheduled maintenance, and — potentially the most disruptive — no warning at all before something shuts down. Taking each in order: with planned maintenance, Azure sends a message ahead of time, our system intercepts it, and we understand a VM needs to be redeployed onto new hardware. We ingest that message, identify the affected nodes, and redeploy gracefully in a way that doesn’t impact or degrade the system. So in a planned scenario, there’s time, and the system handles it accordingly.
With unplanned maintenance, there’s usually still some kind of message beforehand that the system intercepts and uses to redeploy the node in place — if there’s any early warning, which there sometimes is (I believe the SLA is around 15 minutes, though it often happens faster), our software automatically intercepts it. And even if something shuts down without notice at all, the same thing happens — the system automatically recovers. This is what Yitzchak was pointing to earlier: something we tested extensively in the POC and continue to test regularly, and it’s proven out in real production scenarios that we can handle these kinds of node recoveries in stride without affecting the system. I hope that answers the question.
Yitzchak Wahnon: If I can add — the architecture is also planned around things like availability sets, so the chances of losing multiple resources at the same time are almost zero.
Ori Weizman: Great point. Looks like that was the last question, so I’ll leave everyone with this: you can find us on LinkedIn — you can find Yitzchak, and you can find me; my LinkedIn is just my first and last name after linkedin.com/in/. If you search Yitzchak’s name, he’ll come right up too. This was a great conversation — I really enjoyed talking with you, Yitzchak, and I appreciate you being here. I’ll be seeing you again soon, so you won’t get too big a break from me. Thanks so much, everyone, for joining — looking forward to seeing you on the next one. Yitzchak, thanks again.
Yitzchak Wahnon: Okay, good evening, everyone. Thank you.
As a global trading platform handling 2 million transactions a day, eToro depends on consistent application performance to meet demanding SLAs, but keeping workloads fast and reliable in the cloud meant constant tuning and overprovisioning.
In this live conversation, Yitzchak Wahnon, Senior Production DBA at eToro, shares how his team changed that by improving application performance, reducing infrastructure overhead, and keeping demanding trading applications reliably within SLA at a fraction of the cost.
You’ll learn how eToro:
- Keeps demanding trading workloads performing without simply adding more infrastructure
- Improves application performance while maintaining demanding SLAs
- Reduces the tuning and overprovisioning required to keep applications running reliably
Meet the Speakers:
Yitzchak Wahnon
Senior Production DBA, eToro
Ori Weizman
Director, Sales Engineering, Silk

