The PTCi Engineer (PetroTechnical Computing Infrastructure Engineer) is a multidisciplinary systems engineer who architecturally designs, builds, integrates, and manages complex IT infrastructure and operational systems. This professional bridges the gap between physical hardware, virtual/cloud solutions, systems software, storage, networking, and automation, ensuring the organization’s technology platforms are secure, scalable, high-performing, and closely aligned to evolving business needs.
Key Responsibilities:
Infrastructure Design & Implementation: - Architect, deploy, and maintain servers (including Linux - Ubuntu, RedHat/Windows), virtualization, and cloud environments (AWS, Azure, GCP, VMware, Docker).
- Design, build, and automate physical and wireless networks for all enterprise devices and business processes.
- Integrate disparate hardware, software, and networking into a unified, resilient ecosystem.
- Ability to advise and guide on hardware options and purchase.
System Administration & Integration:
- Administer, troubleshoot, and optimize Unix/Linux (Ubuntu, RedHat, RHEL), Windows Server OS, and related user/service configurations.
- Manage network infrastructure (Cisco, Aruba), storage solutions (NFS, NAS, Lustre/HPE ClusterStor), and databases (Oracle, SQL Server).
- Oversee license servers, virtual environments (VMs, containers), and manage updates using automated tools (e.g., xCat, Control-M).
Automation & Scripting:
- Develop and maintain scripts (Python, Bash, PowerShell) to automate operational, configuration, and monitoring tasks.
- Deploy and tune system monitoring tools (Prometheus, Grafana, Redfish API).
Security, Compliance & Disaster Recovery:
- Implement robust security measures (firewalls, IDS/IPS, access controls) and enforce policies.
- Plan and execute system backup, DR, and recovery procedures.
- Ensure system compliance with audits, standards, and policies.
Performance Tuning & Support:
- Continuously monitor infrastructure, manage capacity planning, and optimize throughput (e.g., HPC job scheduling, plot/print services).
- Troubleshoot and resolve complex failures, bottlenecks, and performance anomalies.
- Document workflows and collaborate with internal and third-party support teams.
Skillset Requirements:
Education & Experience:
- Minimum Bachelor of Science in Computer Science, Information Technology, Network engineering, or related field.
- Minimum 5 years of experience in systems, network, or IT infrastructure engineering field.
Technical Skills:
- Operating Systems: Advanced knowledge of Red Hat Enterprise Linux (RHEL), Ubuntu, and Windows Server.
- Experience managing Linux OS update.
- Administration and maintenance of operating systems above.
- Experience with xCAT and Confluent.
- Experience working with Dell and HPE hardware (servers, tape library, storage devices, and GPU).
- Installation and maintenance experience preferred.
- Managing tickets with vendors.
- Procedural firmware update.
- At a minimum, strong ability to advise the procurement team and solutions architect on hardware specifications.
- Virtualization: Expertise in VMware and Hyper-V at a minimum.
- Basic knowledge of cloud technologies: AWS, Azure, and GCP.
- Database administration: Oracle and SQL Server deployment and administration.
- Physical design, backup/recovery, security management, migration.
- Automation & Scripting: Python and PowerShell at a minimum.
- Bash, Ansible, and Terraform are a plus.
- High-Performance Computing (HPC support):
- Design, upgrade, and maintenance.
- Batch/job scheduling, throughput optimization, performance assessment.
- Cluster Management: Expertise in orchestrating compute nodes using job schedulers like Slurm or PBS Pro.
- Storage Systems: Understanding parallel and distributed file systems (e.g., Lustre, GPFS/Spectrum Scale, or Ceph) for high-bandwidth I/O operations.
- Operating Systems: Deep understanding of Linux OS internals, kernel tuning, and system administration specific to compute nodes.
- Networking:
- Virtual Network Management: Administer virtual routers, switches, firewalls, load balancers, VPNs, and VLANs using both CLI and graphical interfaces.
- Implement and enforce network security best practices in virtual environments, covering encryption, authentication, access control, and auditing.
- Experience and familiarity with LAN/WAN, TCP/IP, DNS, DHCP, BGP, OSPF, EIGRP, VPNs, VLANs, NIS, and Samba.
- Experience with Cisco and Aruba.
- Storage infrastructure:
- NFS, NAS, Lustre/ClusterStor, backup/recovery, file system tuning.
- System Monitoring:
- o Prometheus, Grafana, Redfish API.
Other Skills and Abilities:
- Excellent troubleshooting and diagnostic ability.
- Strong workflow documentation and ability to create system/network diagrams.
- Effective communication, time management, teamwork, and vendor coordination.
- Ownership, flexibility, and dependability in multidisciplinary settings.
- Team Player with effective communication skills to interact effectively at all levels
- Creative, enthusiasm and willingness to learn in a fluid and fast-paced environment
- Ability to drive projects forward from a technical standpoint, challenging existing processes and procedures where necessary.
Preferred Qualifications:
- Previous experience managing HPC clusters, upgrade, and maintenance.
- Previous experience managing Dell and HPE servers, and Linux servers.
- Certifications (CCNA/CCNP/CCIE, CompTIA Network+, JNCIE-ENT, WCNA, CWNP, ITIL).
- Data center or large-scale network management (multi-vendor).
- Enterprise monitoring, updating, compliance, and audit experience.