Quick Overview
- CategoryOther Opportunities
- LocationKathmandu
- Job TypeFull Time
- DeadlineAug 09, 2026
Required Skills
Job Description
About the role
We're looking for a Support Engineer who is equal parts data platform administrator and detective. You'll keep our databases, caching and event-streaming systems healthy, be the person who works backwards from a customer symptom to the root cause in the product and step up to keep services running when things go wrong. This is a hands-on role for someone who enjoys the puzzle of "why is this happening?" as much as the discipline of keeping systems well-run — and who stays calm under pressure when a system is down.
What you'll do
- Administer, monitor and tune production and staging databases across MongoDB and PostgreSQL, including backups, restores, replication and failover.
- Manage and troubleshoot Redis (caching, session stores, rate limiting) and Kafka (topics, consumer groups, partitioning, lag), keeping the data layer performant and reliable end to end.
- Respond to server downtime and service outages — including situations where you are the only person available — taking ownership of the incident, restoring service and keeping stakeholders informed until resolution.
- Investigate reported bugs and unexpected behaviour, tracing issues from the customer-facing symptom through logs, traces and message flows to the underlying data and queries.
- Diagnose and resolve performance problems: slow queries, indexing gaps, lock contention, connection pool exhaustion, cache misses, consumer lag and schema-level inefficiencies.
- Write, review and optimise queries and aggregation pipelines to support both live troubleshooting and recurring reporting needs.
- Manage schema changes, migrations and data-integrity checks, working closely with engineering to roll them out safely.
- Reproduce customer issues in test environments, isolate whether the cause is data, configuration, messaging or code, and hand off clear, well-evidenced reports to the product and engineering teams.
- Maintain and improve monitoring, alerting, tracing and runbooks so recurring problems are caught early and resolved consistently.
- document post-incident reviews, capturing root cause and follow-up actions to prevent recurrence.
- Act as an escalation point for complex support tickets that touch the data layer.
What we're looking for
- 1-3 years' experience in support engineering, DevOps, SRE or database administration, ideally in a product or SaaS environment.
- Hands-on experience administering both MongoDB and PostgreSQL in production (or deep experience in one and genuine working knowledge of the other).
- Working experience with Redis and Kafka, and an understanding of how caching and event-driven messaging behave under load and fail.
- Confidence reading and writing SQL, plus MongoDB queries and aggregation pipelines.
- Exposure with distributed tracing and observability, ideally Open Telemetry, to diagnose issues across a distributed system.
- A methodical, evidence-led approach to debugging — comfortable moving between logs, traces, database internals and application behaviour to find root cause.
- Proven ability to stay calm and act decisively during outages, including working independently to restore service when colleagues are unavailable.
- Understanding of indexing, query plans (e.g. EXPLAIN), replication and backup/recovery concepts across both databases.
- Willingness to participate in an on-call rota and respond to incidents outside normal hours.
- Clear written communication: you can turn a messy incident into a concise, reproducible report.
- A customer-focused mindset — you understand that behind every ticket is someone who's blocked.
Nice to have
- Experience with a metrics/visualisation stack (e.g. Prometheus, Grafana, Datadog, Jaeger) alongside Open Telemetry.
- Scripting for automation (Python, Bash or similar).
- Exposure to a microservices architecture and understanding of how issues propagate across services.
- Experience in a SaaS or product support environment.