Key Responsibilities
Develop and maintain data pipelines using PySpark
Process and analyze large-scale datasets in distributed environments
Design and implement ETL/ELT workflows
Optimize Spark jobs for performance and scalability
Work with data stored in HDFS, Hive, or cloud storage (S3, ADLS)
Collaborate with data engineers, analysts, and business teams
Ensure data quality, integrity, and governance
Debug and troubleshoot data processing issues
Automate workflows using scheduling tools (Airflow, Oozie, etc.)
Write clean, scalable, and efficient code
Required Skills & Qualifications
Technical Skills
Strong proficiency in Python and PySpark
Good experience with Apache Spark (RDDs, DataFrames, Spark SQL)
Knowledge of Hadoop ecosystem (HDFS, Hive)
Experience in ETL pipeline development
Familiarity with SQL and database concepts
Experience with data formats (Parquet, ORC, JSON, CSV)
Basic understanding of distributed computing concepts
Exposure to version control tools (Git)
Technology->Big Data - Data Processing->PySpark
Source: Infosys careers — Read the original posting and apply
This role is listed by the employer on its own careers site. Kaam Ki Khoj does not process applications for it.
Upload your CV once — we read it, build your profile and put you in front of every employer hiring on Kaam Ki Khoj. No forms, no fees.