Perks and Benefits
- 100% Paid Healthcare
- 10% 401k in every paycheck
- 100% Fully Vested!
NOTE – Our positions require a Top Secret clearance, as well as the favorable completion of a polygraph. Applicants must be authorized to work in the U.S. We are unable to sponsor an employment Visa.
What You’ll Be Doing (We don’t love the bullet points, but we love the work!)
In this key SWE role, you'll get to support Data Science enablement, Preprocessing, Validation, and Anomaly detection. Preparation of Structured, Semi-structured, and Unstructured datasets for Downstream Analytics and
Machine Learning dataflow design, data transport mechanisms, and Apache Spark based distributed processing. In this role, the you'll be responsible for designing, implementing, and optimizing data ingress/egress pathways to ensure efficient, scalable, and reliable processing of the organization’s analytics workloads.
Required Skills
- Using the Linux CLI and Linux tools
- Developing Bash scripts to automate manual processes
- Recent software development experience using Python and Java
- Experience using Apache Airflow (DAG design, scheduling, operators, sensors) to orchestrate, schedule, and monitor complex workflows
- Familiar with Distributed Big Data processing engines including Apache Spark
- Experience with SQL technologies such as MySQL, MariaDB, and PostgreSQL for querying, joining, and aggregating large datasets
- Experience using Jupyter Notebook
- Experience with data wrangling and preprocessing using tools such as pandas, NumPy
- Experience working with structured, semi-structured, and unstructured data such as Parquet, JSON, CSV, XML
- Familiarity with data quality concepts, data validation, and anomaly detection
- Experience with Git Source Control System
Desired Skills
- Familiar with HPC Job Scheduling tools including Slurm
- Experience using the Atlassian Tool Suite (JIRA, Confluence)
- Maintains a professional relationship with coffee and a recreational interest in witty exchanges. ☕