Work at Tally Solutions

AI systems for product support, software development, and enterprise workflows. The work connects application behaviour with the retrieval, evaluation, and infrastructure needed to make it dependable.

Product support

TARA — Multimodal Support Assistant

Screenshots, audio, and video often carry the details that a written support question leaves out. Specialized extraction agents turn those inputs into focused issue descriptions, preserving error codes and business context for documentation-grounded answers.

Country-based filtering selects relevant source material before generation. Layered safety checks screen inputs, and the support agent keeps trusted instructions separate from external content. Dedicated escalation logic recognizes explicit requests for human help and unresolved issues. Together, these capabilities extend self-service support across different ways of describing a problem.

Region-Aware Website Retrieval

A corporate website assistant uses Vertex AI Search and geographic filtering to retrieve region-specific pricing and compliance information. A shared content base serves localized answers without maintaining separate copies for each market.

Developer tools

Automated Code Review

An automated review platform combines static analysis with a multi-agent workflow to investigate changes in their repository context. Targeted subagents inspect related code and return source locations; the system rereads that evidence before validating findings and attaching them to the pull request.

Harness engineering covers bounded tool execution, context budgets, cancellation of superseded reviews, and recovery from model and scanner failures. Bitbucket and Jira integrations place findings in the existing developer workflow, with tracing that makes execution and model usage visible.

Agentic Code Remediation

SonarQube diagnostics feed an agent that inspects the surrounding code, applies targeted fixes, and opens pull requests for review. The workflow carries routine remediation from issue detection to a reviewable change.

Requirements

Enterprise Agentic SDLC Framework

A requirements-qualification workflow turns fragmented documents and stakeholder input into a working Business Requirements Document. Clarification and review agents identify missing information, ask focused questions, and refine the document. Structural checks and LLM-as-a-judge evaluation assess its completeness and consistency.

Section-level revision keeps the current document coherent as answers arrive. Validated tool actions, persistent checkpoints, and recovery paths preserve progress across interrupted executions, allowing attention to stay on unresolved business requirements.

Evaluation

LLM Evaluation and Regression Testing

Model and configuration changes can alter answer quality even when application logic stays the same. A reusable framework for TARA and TallyPal automates test execution, assessment, and reporting, combining G-Eval completeness scoring with Answer Relevancy.

The approach accounts for valid responses that are more comprehensive than their reference answers and gives ambiguous results a separate review path. Automated assessment reduced manual review volume by approximately 75%, making evaluation more repeatable across application changes.

Multi-Model Evaluation Platform

A self-hosted evaluation platform brings models from multiple providers behind unified routing. Side-by-side assessment makes differences in response quality and inference cost easier to compare when choosing models for a particular workload.

Shared runtime

Unified AI Runtime

A shared backend supports TallyPrime Copilot and a multimodal TDL programming assistant. Copilot combines documentation retrieval with procedural video recommendations; the TDL assistant interprets text, images, audio, and screen recordings to guide programming assistance.

Application-specific agents use a common lifecycle for conversation memory, safety checks, cancellation, persistence, and tracing. Temporary agent sessions and external memory reconstruct context across workers, while reusable execution rules allow new applications to share the same service foundations.

Infrastructure

Multi-GPU LLM Inference

A vLLM service supports internal applications and model experimentation on a four-GPU machine. Model parallelism and quantization accommodate models that exceed a single GPU’s memory, turning the available hardware into a shared inference capability.

LLM Tracing and Observability

Langfuse instrumentation connects model calls, retrieval, and tool execution into inspectable traces. Latency and token-cost breakdowns show where processing time and inference spend accumulate, with prompt versioning and dashboards for comparing text, audio, and video interactions.

Runtime and Deployment Reliability

Reliability work addressed memory pressure in a FastMCP service and stale iframe clients after backend changes. Explicit garbage collection stabilized request handling, while URL-based cache busting refreshed stale clients and resolved the deployment mismatch.

Semantic Caching

Semantic caching design combines vector-based answer reuse with retrieval as a fallback. Data preparation screens personal information locally and checks standalone-query eligibility, with separate review criteria for answer validity and preserved query intent. Embedding benchmarks assess the serving requirements behind the design.

Video Metadata Optimization with TOON

Converting video metadata from JSON to TOON reduced the metadata payload’s token count by approximately 60%, retaining the same fields and content. Manual comparisons and automated relevance evaluation checked recommendation quality across the change.

Enterprise knowledge

TallyPal — Enterprise Knowledge Retrieval

TallyPal brings HR policies, IT support, and product documentation into a shared knowledge layer. Metadata-based retrieval selects information for the employee’s organizational context, helping relevant answers surface across otherwise separate internal systems.

Operational Email Ingestion and Indexing

Automated ingestion and indexing turn historical and incoming operational emails into searchable institutional knowledge. Earlier decisions and working context become available for onboarding and day-to-day questions, reducing dependence on colleagues to reconstruct that history.

Graph-Based Documentation Retrieval

A graph-based retrieval system preserves the headings, tables, and relationships in heterogeneous product documentation. Domain entity extraction narrows the search to relevant content, cross-encoder reranking refines the results, and query decomposition gives compound questions separate retrieval paths.