- Design and run randomized experiments on tier routing, MCP coverage, permission configuration, repository context quality, and budget headroom
- Own the analytical layer of the measurement program, including work classification, evaluation design, longitudinal within-unit analysis, and staggered-adoption estimates
- Build and validate LLM-as-judge and classification pipelines with sampling, hand-labeled ground truth, precision and recall measurement, and revalidation
- Extend AI capabilities across testing, environment and data setup, migration and modernization, code review, security remediation, and certification evidence assembly
- Work directly with constrained teams to identify delivery bottlenecks and target capabilities accordingly
- Build evaluations for internal AI capabilities, including golden sets, regression suites, groundedness, answer-quality scoring, cost, and latency telemetry
- Identify effective practitioner behaviors, document and teach practices, and publish practices rather than rankings
- Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry
- Report findings to engineering leadership and finance, including what is working, delivery constraints, and the limits of each claim
Requirements
- 7+ years spanning software engineering and quantitative analysis
- Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice
- Experimental design and causal inference, including randomized and quasi-experimental designs, difference-in-differences, instrumental variables, and hierarchical models
- Strong Python and SQL, with a statistical stack such as pandas, statsmodels, scikit-learn, or R
- Data collection from operational systems and APIs; robust sampling design
- Identity resolution across systems and complex data joins
- Exploratory analysis, distributions, cohort and time-series analysis, and reporting of coverage and limitations
- Familiarity with the software delivery lifecycle, including code review, CI/CD, test strategy, release and change management
- Git and GitLab at instrumentation depth, including merge request and pipeline data models, diffs and SHAs, merge, squash, rebase, cherry-pick, APIs, and hooks
- Jira and Confluence integration experience, including REST APIs, changelogs, page version history, GitLab–Jira development panel, fields, and labels
- Care with personnel-adjacent data and aggregate reporting
- Communication with executive and engineering audiences
- Preferred MCP servers and clients or comparable connector frameworks
- Agent frameworks such as LangGraph, LangChain, Bedrock Agents, Strands, or equivalent
- Enterprise deployment of coding assistants and telemetry
- Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance
- Confluence and Jira as MCP-connected systems, including permission propagation, scoped credentials, and audit logging
- Evaluation tooling such as Ragas, DeepEval, or Bedrock model evaluation; LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing
- Amazon Bedrock and AWS cost and usage data
- Engineering productivity frameworks such as DORA, DX Core 4, or SPACE
- Program analysis, test generation, or developer tooling research
- dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling
- BI and visualization tooling; summary-table-based reporting
- Queueing and flow analysis, including utilization, batch economics, and constraint identification
Core Competencies
Demonstrates expertise in experimental design and causal inference, with a strong focus on LLM applications and quantitative analysis. Proficient in Python and SQL, with experience in data collection, exploratory analysis, and software delivery lifecycle.
Highest-signal resume keywords
- LLM Application Development
- Experimental Design and Causal Inference
- Python and SQL Proficiency
- Data Collection and Sampling Design
- Git and GitLab Expertise
Hard Skills
- Experimental Design
- Causal Inference
- Python
- SQL
- Data Collection
- Exploratory Analysis
- LLM Applications
- Statistical Analysis
- Sampling Design
- Evaluation Tooling
Soft Skills
- Communication with Executive Audiences
- Documentation and Teaching Practices
Industry Keywords
- Software Engineering
- Quantitative Analysis
- AI Capabilities
- Software Delivery Lifecycle
- Engineering Productivity Frameworks
Tools & Technologies
- Git
- GitLab
- Jira
- Confluence
- Amazon Bedrock
- AWS
- Ragas
- DeepEval
- LangChain
- Airflow
#J-18808-Ljbffr
Senior Software Engineer – Applied AI Arbeitgeber: Jobtailor
Als Front Office Supervisor in unserem dynamischen Team bieten wir Ihnen die Möglichkeit, in einem unterstützenden und freundlichen Arbeitsumfeld zu wachsen. Wir legen großen Wert auf die berufliche Entwicklung unserer Mitarbeiter und bieten regelmäßige Schulungen sowie die Chance, Verantwortung zu übernehmen. Unsere Lage ermöglicht es Ihnen, Teil einer lebendigen Gemeinschaft zu sein, während Sie gleichzeitig die Standards unseres Franchise-Partners einhalten und unseren Gästen einen unvergesslichen Aufenthalt bieten.