Remote Jobs at Alpaca

·

Alpaca

Applications are invited from suitable and qualified candidates for the Following Remote Jobs at Alpaca.

About Alpaca

Alpaca is a financial technology company that provides API-based brokerage and trading infrastructure for individuals, developers, fintech companies, and financial institutions. Headquartered in Silicon Valley, the company develops technology that allows businesses and developers to build investing applications, brokerage platforms, and algorithmic trading systems using its APIs. Alpaca’s services provide access to financial markets and support trading across assets such as U.S. stocks, exchange-traded funds (ETFs), options, and cryptocurrencies, while its market data tools provide access to real-time and historical financial information. Its Broker API is designed to help fintech companies and institutions build end-to-end brokerage experiences, including account opening, funding, trading, and portfolio management, while the Trading API enables retail, algorithmic, hedge fund, and proprietary traders to develop automated trading strategies.

Alpaca also emphasizes security, regulatory compliance, and developer-friendly technology, with brokerage services provided through regulated entities and infrastructure designed for global financial applications. The company works with fintechs and financial institutions in multiple countries and focuses on expanding access to financial markets through modern technology and API-driven solutions. With its combination of financial technology, brokerage infrastructure, market data, and developer tools, Alpaca provides career opportunities across software engineering, fintech, product development, cybersecurity, finance, operations, customer support, and other professional fields.

Summary

  • Company: Alpaca
  • Job Opening: 8 Positions
  • Job Type: Full Time
  • Eligible Countries: All Countries
  • Location: Remote

backpack

Job Opening: 8 Positions

See More Posts in Jobs, Scholarships, Technology, Career/Motivations, Football News Feeds

Join Job Whatsapp Channel, Scholarship Whatsapp Channel, Tech Whatsapp Channel, Follow Our Twitter (X) Channel.

1. Remote Tech Jobs at Alpaca – 4 Positions

Click Here for Details and Apply

2. Job Title: Senior Site Reliability Engineer

  • Job Type: Full Time
  • Eligible Countries: All Countries
  • Location: Remote, USA

Description:

As a Site Reliability Engineer at Alpaca, you’ll help keep our brokerage platform reliable, observable, and operable as we grow – working across our cloud infrastructure, Kubernetes platform, observability stack, messaging layer, and data layer. We’re especially interested in candidates with strong PostgreSQL fundamentals who’d like to grow into deeper ownership of our database reliability posture: PostgreSQL sits on the trading-critical path, and we want this person to spend a meaningful share of their time leveling it up while still being a well-rounded SRE the rest of the week.

Things You Get To Do

  • Operate production day-to-day – oncall, incident response, postmortems, and the follow-ups that actually close the loop.
  • Own reliability practice – define and refine SLIs/SLOs and error budgets, and help product teams live within them.
  • Strengthen our observability across metrics, logs, traces, and alerting.
  • Ship infrastructure through code in a GitOps workflow – cloud resources and Kubernetes workloads alike.
  • Look after PostgreSQL: performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines.
  • Mentor engineers on reliability and database fundamentals through code review, design review, and pairing.

See Also:

Requirements:

  • 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.
  • Hands-on experience operating production services on Kubernetes, and shipping infrastructure as code in a GitOps workflow.
  • Solid working knowledge of PostgreSQL in production — query plans, pg_stat_*, indexing and schema trade-offs, and what a safe online migration looks like on a non-trivial table.
  • Cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and comfort debugging cross-service connectivity.
  • Comfortable with a modern observability stack and proficient with Linux at the operator level.
  • Practiced in incident response – calm under pressure, structured debugging, postmortems that drive change.
  • At least working proficiency in Go or Python, plus strong written and verbal communication.
  • Genuine interest in databases and in growing your PostgreSQL/DBA expertise.

Who You Might Be (Nice-to-Haves):

  • Deeper PostgreSQL experience: large clusters at OLTP load, online migrations on big tables, HA/DR ownership, connection pooling at scale, or change-data-capture pipelines.
  • Experience with typed SQL access layers in Go (e.g. pgx, gorm, sqlc).
  • Production experience with messaging systems at scale (e.g. RabbitMQ, Kafka, Redpanda).
  • Security & compliance experience in a regulated environment (SOC 2, secrets management, audit logging).
  • Familiarity with trading, brokerage, or other regulated fintech domains.

How We Take Care of You:

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Click Here to Apply

3. Job Title: Incident Operations Lead (EMEA/AMER)

  • Job Type: Full Time
  • Eligible Countries: All Countries
  • Location: Remote, USA

Description:

Lead the team that commands Alpaca’s most critical incidents. You will build the function and then keep raising its bar: the severity model, the escalation and communication paths, 24×7 follow-the-sun coverage, and the KPIs that prove it is improving. You will do that across boundaries – with the engineering teams who own the services, with SRE on reliability standards and on-call readiness, with Risk on financial and regulatory materiality, with our partner communications teams on what reaches a customer, and with reliability programme management on what happens after.

You will own how well we respond. Not the fix, not the partner communication, and not the reliability standard. Holding that line is a deliberate part of the design and a core part of the job.

Things You Get To Do

  • Build the team and stand up 24×7 command. Recruit and certify Incident Commanders, build a follow-the-sun rotation across APAC, EMEA and AMER with warm handoffs at every regional boundary, and carry a rostered slot yourself. Keep the team sharp between real incidents with game days, tabletop exercises and simulations, and coach them through the live ones. Build a blameless review culture that treats an outlier as a process gap rather than a person’s failure.
  • Own the process, and keep raising it. Drive severity maturity with Risk on financial and regulatory materiality – in a regulated brokerage a severity call can also start a reporting clock, so the model has to map cleanly onto those thresholds. Own the escalation path and what happens when a page goes unanswered, agree the thresholds for taking an incident to engineering leadership, and keep the service catalogue and its ownership current – time spent working out who owns a failing service is customer impact.
  • Own both bridges. Your team connects the engineers and technical support fixing the problem, who need uninterrupted focus, to the partner communications teams, who need a continuous and accurate feed. You open the channel, supply the facts and hold the update cadence to account. Afterwards, your team runs the retrospective with SRE – who own the technical depth – and builds the post-incident package while the room is still warm, every item ticketed, owned and tagged, delivered inside a service level you define and then hold, before reliability programme management drives it to closure. You then synthesise the discussion into short, digestible learnings and publish them to the whole engineering organisation, so one team’s failure becomes everyone’s lesson instead of a document three people read.
  • Own the KPIs. Time to respond and time to mitigate end to end, including the definitions and data hygiene beneath them: what separates mitigated from resolved, and whether a timestamp means what it claims. Establish a defensible baseline before committing to targets, then move them by severity. Review and approval, not authorship, is where postmortems stall, so report overdue reviews by team and incident with a next action against each.
  • Build it as a product, then automate it with AI. Everything is documented, versioned and deployable, so you can stand up command from the artefacts alone; the process needs to scale considerably faster than the team. The automation we are after is AI workflows and agents rather than scripts and dashboards – agents that set an incident up, assemble the timeline as it runs, draft the RCA and the action package, and chase the update that is due or the review that is overdue. You own that roadmap: what an agent may do unsupervised, what still needs a commander’s judgement, and the decision-tree quality that makes either of them safe.

Requirements:

  • You have stood up an incident command or major-incident function, not only worked inside one – you have owned the severity model, built the roster and driven adoption across teams.
  • 5+ years in production engineering, SRE or technical operations, including hands-on command of high-severity incidents.
  • You have led a distributed team across time zones and run a 24×7 rotation.
  • You get engineers you do not manage to do things, and you can defend a severity call to someone who disagrees with it.
  • You have built reliability metrics people trust, and you know the difference between improving a number and improving reality.
  • You are disciplined about scope. You can say “that is not ours” and route it, in the middle of an outage, without leaving a gap.
  • You write well enough that your process documents actually get used, you can hold a bridge calm under pressure, and you can brief an executive mid-incident without either downplaying it or dramatising it.
  • You understand FinTech and the trust stakes of API-driven financial platforms.
  • You use AI and agentic automation to remove toil rather than to add tooling.

Who You Might Be (Nice-to-Haves)

  • Formal incident command training – ITIL, Major Incident Management or crisis management.
  • You have run a certification, game day or drill programme, or built a pool of certified responders beyond your own headcount.
  • Experience with modern incident management and on-call platforms.
  • You have built a service catalogue or ownership registry that people actually maintained.
  • You have worked with programme management or reliability functions to convert incident follow-ups into funded roadmap work.
  • Familiarity with incident reporting obligations in regulated financial services – DORA, Reg SCI, FINRA or equivalent.
  • Online securities trading or capital markets experience, or another regulated, market-hours-sensitive domain.
  • You have deployed the same operating model into a second region or entity.

How We Take Care of You:

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Click Here to Apply

4. Job Title: Incident Operations Commander

  • Job Type: Full Time
  • Eligible Countries: All Countries
  • Location: Remote, USA

Description:

Serve as the on-duty commander for Alpaca’s most critical incidents, directing cross-functional response to restore service quickly, keeping the right people engaged and informed, and making sure every incident leaves behind something the organisation can act on.

You do not fix the outage. You make the response reliable: correct severity, the right engineers in the room, mitigation that does not stall, leaders informed in time, and follow-up work that survives the call.

Things You Get To Do

  • Command incidents end to end. Take command from declaration to mitigation, keeping responders focused on stopping customer and partner impact as fast as possible. Run the bridge, keep observers out of the responders’ way, and name a stall out loud when you see one.
  • Classify and hold the line on severity. Set severity at declaration and re-check it as facts arrive. Risk advises on financial and regulatory materiality; the call is yours.
  • Engage the right people, fast. Identify the owning team by service, symptom and blast radius, page them, and expand the responder set the moment the first team is wrong or not enough. When a page goes unanswered, escalate – and escalate the escalation. Bring in the leaders who must make business calls: feature flags, traffic shedding, failover, freeze-or-ship.
  • Hold the bridge, and protect the people fixing it. Keep engineering and technical support uninterrupted – questions from stakeholders, partners and executives come to you. Be the single source of truth to the partner communications team on impact, severity and timing: you decide when a status page update or partner contact is needed, they write and send it, and chasing a late or stale update is yours.
  • Run follow-the-sun handoffs. Deliver warm, high-fidelity handoffs across regions: current impact and severity, mitigation path and next actions, who is in the room, outstanding decisions, and what must not be dropped. The incoming commander confirms ownership before you step away – command never goes dark at a region boundary.
  • Close the loop, on the clock. Maintain the timeline of facts as the incident runs rather than reconstructing it afterwards – in a regulated business that record has to hold up long after the call ends. Once mitigated, make sure a blameless retrospective is scheduled with a named owner and a timebox, and record where the cause sits – that choice sets which follow-up items are mandatory. Every action item needs a real ticket, one named accountable, a priority and a category, delivered inside the agreed service level. If a postmortem produces nothing but low-priority items, treat that as a signal the analysis stopped early and escalate to SRE rather than passing it on.
  • Automate the coordination away. Coordination is the part of this job that should eventually belong to a machine. Every manual prompt you send – the update that is due, the question nobody answered, the partner nobody contacted – is a candidate for automation, and the direction we are heading is AI handling the routine so commanders can spend their attention on judgement. You get us there by working to the decision trees, saying where they are wrong, and being honest about which of your instincts are actually rules.

Requirements:

  • 4+ years commanding or co-commanding high-severity incidents in a production engineering, SRE or technical operations environment.
  • You direct technical responders under pressure without being the person writing the fix.
  • You make and defend crisp severity and escalation decisions, and you take charge without waiting to be asked. Command means waking senior people at 03:00, interrupting an executive, and telling an experienced engineer to stop what they are doing – with an audience watching. It is a visible, directive role and it needs to be instinctive.
  • You can read a dashboard and judge for yourself whether impact has actually stopped.
  • You communicate clearly with engineers, executives and partner-facing stakeholders – and you know the difference between briefing the comms function and speaking for the company.
  • You are comfortable holding other teams to account in the moment, across a reporting line that is not yours, without turning it into friction.
  • You thrive in a follow-the-sun model with clean cross-region handoffs.
  • You understand FinTech concepts and the trust stakes of API-driven financial platforms.
  • You use AI tools and agentic automation to reduce manual toil and speed up response.
  • You will work a regional coverage window as part of a global 24×7 Incident Commander roster.

Who You Might Be (Nice-to-Haves)

  • Formal incident command training – ITIL, Major Incident Management or crisis management.
  • Experience with modern incident management and on-call platforms.
  • You have written severity rubrics, decision trees, escalation matrices, runbooks or incident playbooks.
  • You have commanded in game days, tabletop exercises or incident simulations, not only in production.
  • You have partnered with problem management or reliability programme functions to roadmap incident follow-ups.
  • Online securities trading or capital markets experience, or another regulated, market-hours-sensitive domain.

How We Take Care of You:

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Click Here to Apply

Deadline

Not Specified

Click Here to See other Jobs.

Get a professional, ATS compliant CV, and Cover Letter from an Expert.

(See tips on how to write a professional CV and a sample cover letter.)

Important: See Helpful Career Resources

See More Posts In:

, , , , , , , , , , , , ,

Share Post to:

Subscribe to Get Notifications: